All posts

October 5, 2026 · Temple Compute

Generative Ligand Design with DrugFlow and Boltz-2

Virtual screening answers one question: which of the molecules I already have binds this pocket best. Structure-based generative ligand design answers a different one: what molecule should exist for this pocket. You do not hand it a library. You hand it a protein and a reference ligand that marks the site, and it proposes molecules built for that site.

The catch is that a generative model proposes. It does not score. A model trained to put plausible atoms in a pocket will happily hand you ten molecules with no ranking between them, and "plausible" is not "binds". So the useful unit of work is not the generator alone. It is a loop: generate in the pocket, then score each proposal with a separate affinity model, then rank.

Our new DrugFlow + Boltz-2 affinity workflow is that loop, as six stages you can run end to end.

Generate, then score, then rank

Generate. DrugFlow takes the target protein as a PDB file and a reference ligand as an SDF. The reference ligand's binding site defines the pocket; a distance cutoff around it, 8.0 ångström by default, defines how much of the protein the model conditions on. It then samples novel molecules for that pocket and writes them as an SDF. You set how many with --n_samples, and you can cap molecule size or ask the model to resample until its quality filters pass.

Prepare. The scoring model needs a sequence and a SMILES string, not an SDF. This stage reads the generated SDF with RDKit, canonicalizes each molecule to SMILES, extracts the longest protein chain's sequence from the PDB with Biopython, and writes one Boltz affinity YAML per molecule into a single folder. It also writes a {molecule_name: canonical_smiles} JSON map.

One detail matters here: each YAML is named after its molecule. Boltz derives the prediction name from the YAML stem, so mol_3.yaml produces affinity_mol_3.json, and that name is the key the ranking stage joins against the SMILES map. Name every file mol.yaml and the join has nothing to work with.

Score. Boltz-2 runs once over the whole input folder and writes one affinity JSON per molecule. This is the same affinity model behind our Boltz-2 virtual screening workflow. The difference is only where the ligands came from.

Rank. The last stage walks the predictions, reads each affinity_pred_value, converts it to an approximate binding free energy, joins the SMILES back in, and writes a CSV sorted best first. Boltz-2 reports that value as roughly log10(IC50) with IC50 in µM, so treating IC50 as a Kd proxy at 298 K gives:

ΔG ≈ 1.364 · (affinity_pred_value − 6)   kcal/mol

That is an approximation, and worth being blunt about: it is good for ordering candidates against each other, not for quoting a free energy. The output table carries affinity_pred_value and the binding probability alongside ΔG so you can look at the raw numbers.

If you want a cheaper first pass before any of this, empirical docking is still the fastest filter, see the AutoDock Vina docking workflow.

Why this needs GPU stages and CPU stages

Six stages, three kinds of machine.

Two of them just fetch things. One downloads the DrugFlow checkpoint, about 170 MB, from Zenodo. The other clones DrugFlow at a pinned commit, because DrugFlow is not published as a package: an environment can supply its dependencies but not its code. Each needs one core and a network route, and each is skipped once its output is already there. The download writes to a .part file and moves it into place only on success, because a truncated file sitting at the final path would look complete forever and be reused as a corrupt checkpoint on every later run.

Generation wants a GPU. It is the stage that sets the hardware bill.

Preparation, prediction and ranking are CPU work. Preparation is seconds of RDKit and Biopython. Prediction defaults to a CPU accelerator so it runs anywhere, though it is slow there. Ranking is stdlib Python over a folder of JSON.

Pin all six to one GPU machine and you hold a GPU while a sequence gets parsed and a CSV gets sorted. Split them by hand and you are copying archives between hosts, which is where files get lost and stages get run on stale input.

Why containers, and one .sif per source

The dependency sets here do not merge. DrugFlow's own environment pins pytorch=2.2.1, pyg=2.5.1 and ProDy=2.4.0. The preparation stage wants RDKit and Biopython. The prediction stage wants boltz. The ranking stage wants nothing at all beyond a bare python3. There is no single environment that satisfies that set, and on Apple Silicon there is no solve for the DrugFlow half at all, those three pins have no osx-arm64 builds.

So the heavy stage runs in a container, and on a cluster that means Singularity or Apptainer rather than Docker: GPUs live on compute nodes, and compute nodes generally do not give you a Docker daemon. You build one .sif per source, once, from the published DrugFlow image, and the stage runs inside it under an sbatch job.

Two things about that are easy to get wrong.

The image supplies dependencies, not code. generate.py comes from the separate checkout stage, so that checkout has to be visible inside the container. Horus auto-binds every input and output artifact's parent directory at its own path, which covers it. And because Singularity does no path translation, a bound host path appears inside the container at the same location, so every ${...} placeholder in the command is valid on both sides and needs no rewriting.

The image path is site-specific. It is declared in the workflow's artifacts: block as drugflow_sif, next to the protein and the reference ligand, and the task reads it as ${drugflow_sif}. The default is relative, so it resolves in the run directory; point it at an absolute path to share one build across runs. Every path an operator has to supply lives in one block instead of being buried inside a task.

How Horus runs it

Each stage declares its own executor and target. Generation gets the singularity executor on a slurm target with gres: gpu:1 and a time limit. Preparation and prediction each get their own uv environment, built per stage, so RDKit and boltz never have to coexist. Ranking runs on the target's python3. Everything but generation runs locally.

Stages are wired by artifacts, not by order: the checkpoint and the source tree feed generation, generation's SDF feeds preparation, preparation's folder feeds prediction, and prediction plus the SMILES map feed ranking. Horus moves the data across each boundary. You write no copy commands.

To move a stage somewhere else, you change its executor: field. The runtime.command string does not change. That is the same portability rule as every other workflow in the library, and it is why the same YAML file runs on a laptop and on a cluster.

Each stage is also independently resumable. A failed prediction does not cost you the generation run.

FAQ

What is structure-based generative ligand design? It is de-novo molecule generation conditioned on a binding site. Instead of scoring an existing library against a pocket, a generative model proposes new molecules shaped for that pocket, using the protein structure and a reference ligand that marks the site. Scoring is a separate step.

How do I score DrugFlow-generated molecules? Convert each generated molecule to SMILES, pair it with the target sequence, and run an affinity model over the pairs. In this workflow that is one Boltz-2 affinity YAML per molecule and a single boltz predict call over the folder, then a conversion from the predicted affinity value to an approximate ΔG for ranking.

Do I need a GPU to run a Boltz-2 affinity prediction pipeline? Not strictly. The prediction stage defaults to a CPU accelerator so it runs anywhere, and it is slow there. The generation stage defaults to a GPU, and can be switched to CPU by setting the executor's nv: false and --device cpu together, the two must agree, or the container either cannot see the driver stack or does not ask for it.

Why Singularity instead of Docker? Because the GPU is on the cluster and the cluster does not run a Docker daemon. The .sif is built once ahead of time; the executor has no build or pull step, so the image has to exist on the target already. On a machine with Docker, a docker executor is the simpler route.

Does the prediction stage need internet access? Yes, as configured. It passes --use_msa_server, which uses the public ColabFold server, so whichever machine runs that stage needs outbound network access and is subject to that server's rate limits. On a cluster without outbound access, that is the stage that will fail.

Try it

The workflow is open source, with a bundled KRAS example you can run as-is: DrugFlow + Boltz-2 affinity. To run it without building the cluster side yourself, open it in Temple Compute OS and point the generation stage at the hardware you have.

Open Horus