Lock the positions of a scaffold that must survive, grow the rest with DrugEx, then narrow the result stage by stage. Every compound keeps the identifier it was given at generation, so each number on it can be traced back.
Nothing survived the chain. Loosen a range, lower the minimum SA, or uncheck a filter and run it again — the funnel above shows which stage removed the compounds.
Khanh's own round is 100 molecules at 0.2
(lead_generation_until_N.sh). SA is normalised 0–1, not the raw
1–10 Ertl scale, and higher means easier to make. Measured on this scaffold:
0.2 kept 96 of 100, 0.4 kept 28 of 40, and 0.6 and 0.8 kept nothing at all.
Rejects on structure alone, so it has nothing to tune. Compounds matching a known interference or unwanted-chemistry pattern are removed.
Keeps compounds whose predicted pKa falls inside this window. Khanh's scripts use 4 to 6.
A prepared PDBQT, which is what Vina docks against. Path is relative to the project. Choosing a target fills this in.
Khanh's scripts disagree: run_docking_pipeline.sh
uses −5.0, run_filter_docking.sh uses −9.0. More negative is better,
so −9.0 is far stricter. Pick deliberately.
Scores are not reproducible. vinascreen.py passes
no seed to Vina and offers no way to, so the same compound scores differently on
every run — measured spread up to 0.26 kcal/mol against the shipped example. A
compound whose score sits near the cutoff survives or dies by the run.
Question 8 for Khanh.
The sequence in the repository belongs to a third protein, neither FUT8 nor the shipped receptor, so there is no safe default.
One GPU prediction per compound. Measured on dichtator: about 115 s each including the shared MSA, so 25 compounds is roughly 45 minutes and 10,000 would be a fortnight.
This machine has about 7.5 GB of VRAM and a protein of 369 residues needs about 17 GB, so affinity runs elsewhere. A venv binary is not on a non-interactive ssh PATH, which is why the command is a full path. Khanh’s script hardcodes GPU 1; on dichtator that is the busy card.