October 5, 2026 · Temple Compute
Generative Ligand Design with DrugFlow and Boltz-2
Virtual screening answers one question: which of the molecules I already have binds this pocket best. Structure-based generative ligand design answers a different one: what molecule should exist for this pocket. You do not hand it a library. You hand it a protein and a reference ligand that marks the site, and it proposes molecules built for that site.
The catch is that a generative model proposes. It does not score. A model trained to put plausible atoms in a pocket will happily hand you ten molecules with no ranking between them, and "plausible" is not "binds". So the useful unit of work is not the generator alone. It is a loop: generate in the pocket, then score each proposal with a separate affinity model, then rank.
Our new DrugFlow + Boltz-2 affinity workflow is that loop, as six stages you can run end to end.
Generate, then score, then rank
Generate. DrugFlow takes the
target protein as a PDB file and a reference ligand as an SDF. The reference
ligand's binding site defines the pocket; a distance cutoff around it, 8.0
ångström by default, defines how much of the protein the model conditions on.
It then samples novel molecules for that pocket and writes them as an SDF. You
set how many with --n_samples, and you can cap molecule size or ask the model
to resample until its quality filters pass.
Prepare. The scoring model needs a sequence and a SMILES string, not an
SDF. This stage reads the generated SDF with RDKit, canonicalizes each molecule
to SMILES, extracts the longest protein chain's sequence from the PDB with
Biopython, and writes one Boltz affinity YAML per molecule into a single
folder. It also writes a {molecule_name: canonical_smiles} JSON map.
One detail matters here: each YAML is named after its molecule. Boltz derives
the prediction name from the YAML stem, so mol_3.yaml produces
affinity_mol_3.json, and that name is the key the ranking stage joins against
the SMILES map. Name every file mol.yaml and the join has nothing to work
with.
Score. Boltz-2 runs once over the whole input folder and writes one affinity JSON per molecule. This is the same affinity model behind our Boltz-2 virtual screening workflow. The difference is only where the ligands came from.
Rank. The last stage walks the predictions, reads each
affinity_pred_value, converts it to an approximate binding free energy, joins
the SMILES back in, and writes a CSV sorted best first. Boltz-2 reports that
value as roughly log10(IC50) with IC50 in µM, so treating IC50 as a Kd proxy at
298 K gives:
ΔG ≈ 1.364 · (affinity_pred_value − 6) kcal/mol
That is an approximation, and worth being blunt about: it is good for ordering
candidates against each other, not for quoting a free energy. The output table
carries affinity_pred_value and the binding probability alongside ΔG so you
can look at the raw numbers.
If you want a cheaper first pass before any of this, empirical docking is still the fastest filter, see the AutoDock Vina docking workflow.
Why this needs GPU stages and CPU stages
Six stages, three kinds of machine.
Two of them just fetch things. One downloads the DrugFlow checkpoint, about
170 MB, from Zenodo. The other clones DrugFlow at a pinned commit, because
DrugFlow is not published as a package: an environment can supply its
dependencies but not its code. Each needs one core and a network route, and
each is skipped once its output is already there. The download writes to a
.part file and moves it into place only on success, because a truncated file
sitting at the final path would look complete forever and be reused as a
corrupt checkpoint on every later run.
Generation wants a GPU. It is the stage that sets the hardware bill.
Preparation, prediction and ranking are CPU work. Preparation is seconds of RDKit and Biopython. Prediction defaults to a CPU accelerator so it runs anywhere, though it is slow there. Ranking is stdlib Python over a folder of JSON.
Pin all six to one GPU machine and you hold a GPU while a sequence gets parsed and a CSV gets sorted. Split them by hand and you are copying archives between hosts, which is where files get lost and stages get run on stale input.
Why containers, and one .sif per source
The dependency sets here do not merge. DrugFlow's own environment pins
pytorch=2.2.1, pyg=2.5.1 and ProDy=2.4.0. The preparation stage wants
RDKit and Biopython. The prediction stage wants boltz. The ranking stage
wants nothing at all beyond a bare python3. There is no single environment
that satisfies that set, and on Apple Silicon there is no solve for the
DrugFlow half at all, those three pins have no osx-arm64 builds.
So the heavy stage runs in a container, and on a cluster that means Singularity
or Apptainer rather than Docker: GPUs live on compute nodes, and compute nodes
generally do not give you a Docker daemon. You build one .sif per source,
once, from the published DrugFlow image, and the stage runs inside it under an
sbatch job.
Two things about that are easy to get wrong.
The image supplies dependencies, not code. generate.py comes from the
separate checkout stage, so that checkout has to be visible inside the
container. Horus auto-binds every input and output artifact's parent directory
at its own path, which covers it. And because Singularity does no path
translation, a bound host path appears inside the container at the same
location, so every ${...} placeholder in the command is valid on both sides
and needs no rewriting.
The image path is site-specific. It is declared in the workflow's artifacts:
block as drugflow_sif, next to the protein and the reference ligand, and the
task reads it as ${drugflow_sif}. The default is relative, so it resolves in
the run directory; point it at an absolute path to share one build across runs.
Every path an operator has to supply lives in one block instead of being buried
inside a task.
How Horus runs it
Each stage declares its own executor and target. Generation gets the
singularity executor on a slurm target with gres: gpu:1 and a time limit.
Preparation and prediction each get their own uv environment, built per
stage, so RDKit and boltz never have to coexist. Ranking runs on the target's
python3. Everything but generation runs locally.
Stages are wired by artifacts, not by order: the checkpoint and the source tree feed generation, generation's SDF feeds preparation, preparation's folder feeds prediction, and prediction plus the SMILES map feed ranking. Horus moves the data across each boundary. You write no copy commands.
To move a stage somewhere else, you change its executor: field. The
runtime.command string does not change. That is the same portability rule as
every other workflow in the library, and it is why the same YAML
file runs on a laptop and on a cluster.
Each stage is also independently resumable. A failed prediction does not cost you the generation run.
FAQ
What is structure-based generative ligand design? It is de-novo molecule generation conditioned on a binding site. Instead of scoring an existing library against a pocket, a generative model proposes new molecules shaped for that pocket, using the protein structure and a reference ligand that marks the site. Scoring is a separate step.
How do I score DrugFlow-generated molecules?
Convert each generated molecule to SMILES, pair it with the target sequence,
and run an affinity model over the pairs. In this workflow that is one Boltz-2
affinity YAML per molecule and a single boltz predict call over the folder,
then a conversion from the predicted affinity value to an approximate ΔG for
ranking.
Do I need a GPU to run a Boltz-2 affinity prediction pipeline?
Not strictly. The prediction stage defaults to a CPU accelerator so it runs
anywhere, and it is slow there. The generation stage defaults to a GPU, and can
be switched to CPU by setting the executor's nv: false and --device cpu
together, the two must agree, or the container either cannot see the driver
stack or does not ask for it.
Why Singularity instead of Docker?
Because the GPU is on the cluster and the cluster does not run a Docker daemon.
The .sif is built once ahead of time; the executor has no build or pull step,
so the image has to exist on the target already. On a machine with Docker, a
docker executor is the simpler route.
Does the prediction stage need internet access?
Yes, as configured. It passes --use_msa_server, which uses the public
ColabFold server, so whichever machine runs that stage needs outbound network
access and is subject to that server's rate limits. On a cluster without
outbound access, that is the stage that will fail.
Try it
The workflow is open source, with a bundled KRAS example you can run as-is: DrugFlow + Boltz-2 affinity. To run it without building the cluster side yourself, open it in Temple Compute OS and point the generation stage at the hardware you have.