Drug discovery
Lacuna + AutoDock Vina: Cryptic Pocket Docking
Find cryptic pockets with Lacuna, then dock a ligand library into a ranked site with AutoDock Vina. Meeko and OpenBabel prep, one Vina run per ligand.
What this workflow does
This workflow takes one unbound (apo) receptor structure and a ligand library. It finds the receptor's pockets with Lacuna, including cryptic pockets that are closed or shallow in the input. It then docks every ligand into a ranked pocket with AutoDock Vina. It returns a ranked pocket report and ranked binding energy tables in CSV format.
Lacuna builds a conformational ensemble from the input structure. It detects pockets in every conformer and clusters them across the ensemble. A site that opens in only a few conformers is still reported. A fitted model ranks the clusters and scores how much each site opens relative to the input (its crypticity).
The workflow runs Lacuna with --detector surface-fusion. This pools the
geometric detector with a learned surface detector. The geometric detector only
proposes a pocket where the surface is already concave. A cryptic site violates
that assumption. On held-out CryptoBench data, pooling both detectors takes
coverage from 68.5% to 86.4% and top-five recovery from 57.1% to 73.9%.
The prep stage uses OpenBabel for the receptor and Meeko for the ligands. RDKit embeds each SMILES to 3D with ETKDG and MMFF first. AutoDock Vina then docks each ligand into the search box that Lacuna exported for the chosen pocket.
For pocket discovery without docking, see Lacuna Cryptic Pocket Discovery. For docking into a box you define by hand, see AutoDock Vina Docking.
The compute problem
The pipeline has two halves with different needs.
The discovery half generates the ensemble and runs detection on every conformer.
This is the expensive part of discovery. It needs only lacuna-pockets, a plain
pip package. The default normal-mode backend runs on a CPU.
The docking half needs OpenBabel, Meeko, RDKit, and Vina. These come from
conda-forge. The vina package on PyPI has no macOS arm64 wheel. The conda-forge
build ships Python bindings for osx-arm64, osx-64, linux-64, and linux-aarch64.
One environment for both halves means one dependency solver for two unrelated
stacks.
The dock stage scales with the library size. A 100-ligand library is 100 independent Vina runs. Run them one after another and the wall-clock time grows with every ligand you add.
Ensemble generation is also the stage you want to rerun least. If changing the number of exported pockets meant regenerating the ensemble, every small change would cost the full discovery time.
How Horus solves it
Horus assigns an executor to each stage. The discover and dock_prep stages
run in a uv-managed venv with lacuna-pockets. The prep and dock stages run
in conda-forge environments with the docking stack. The two stacks never share a
solver.
The dock stage is a horus_map task. It runs one Vina docking per ligand
PDBQT, concurrently. Docking time scales with concurrency, not with the ligand
count. Every clone gets a copy of the shared receptor and box.
Discovery and docking-input export are split into two tasks. You change --top
or --format on dock_prep without rerunning ensemble generation and detection.
To move a stage to a cluster, give it a target: and resources: block. Give
discover one to scale ensemble generation. Give dock one and every clone
inherits it. The runtime.command strings do not change. Horus moves the data
across each boundary for you.
A ligand that fails to dock does not stop the run. Its clone records
status: failed in status.json. The summary stage skips it instead of scoring
it as zero.
Pipeline
discover (uv, CPU) receptor.pdb ──► ensemble + detection + ranking ──► pockets/
│ pocket_report.json, pocket_*.pdb
dock_prep (uv, CPU) top 5 pockets ──► docking_inputs/
│ pocket_*_vina.conf, pocket_*_constraint.yaml, pocket_*_site.pdb
prep (conda, CPU) receptor + ligands.smi + pocket_0_vina.conf
│ OpenBabel → receptor.pdbqt, box.json; Meeko → ligands_pdbqt/*.pdbqt
dock[00..NN] (conda) one AutoDock Vina run per ligand ──► docked/<i>/
│ {ligand}_out.pdbqt, status.json
summary (CPU) parse affinities ──► scores.csv + poses.csv
report (CPU) pocket_report.json ──► report.html + pockets.csv
Inputs and outputs
Inputs
receptor.pdb: the unbound target structure. Any single PDB or mmCIF file works. The bundled example is PDB 4LDJ, GDP-bound apo KRAS(G12C) with no switch-II inhibitor present.ligands.smi: oneSMILES [name]per line, or a 3D.sdffile. The bundled set is three demo compounds (imatinib, caffeine, aspirin). It proves the docking mechanics. It is not a curated KRAS library.
Outputs land in results/:
pockets/pocket_report.json: every surviving pocket cluster with rank, druggability, persistence, crypticity, and contact residues.report.html: a ranked pocket table with a Mol* view of the receptor.pockets.csvholds the same ranking as a flat table.docking_inputs/: Vina box configs, Boltz-2 YAML constraints, and pocket PDBs for the top 5 pockets.receptor_box/andligands_pdbqt/: the receptor PDBQT, the resolved box, and one PDBQT per ligand.docked/<i>/: the poses and astatus.jsonfor each ligand.scores.csv: one row per ligand, ranked best first:rank, ligand, best_affinity_kcal_mol, mean_affinity_kcal_mol, num_poses.poses.csv: one row per pose:ligand, pose, affinity_kcal_mol, rmsd_lb, rmsd_ub.
A lower affinity value means stronger predicted binding.
On 4LDJ, Lacuna ranks the GDP pocket first. It surfaces the switch-II cryptic pocket twice in the top ten. KRAS is not in the data Lacuna's ranker was fitted on.
Run the workflow
The prep and dock stages build conda environments from conda-forge. A conda
family tool must be on your PATH. The workflow.yaml file defaults to
micromamba. Set it to mamba or conda if that is what you have.
Install the horus-runtime and the plugins one time:
uv sync
If you do not have uv, install it first:
curl -LsSf https://astral.sh/uv/install.sh | sh
You can also install the packages with pip:
pip install horus-runtime horus-environments
Then run the workflow:
uv run horus run workflow.yaml
To pick the pocket, edit --box-config in the prep stage. The default is
pocket_0_vina.conf, Lacuna's rank-1 site. On 4LDJ that is the GDP pocket.
Swap in pocket_3_vina.conf to dock into the switch-II cryptic pocket instead.
The prep stage keeps only ATOM records, so the bound GDP and waters are
removed before docking.
To change discovery, edit the discover stage. Set --conformers for the
ensemble size and --rank-by for the ranking strategy. Use --min-druggability,
--min-persistence, or --min-crypticity to filter pockets.
To set the search effort, edit the args: of the dock stage. The defaults are
--exhaustiveness 8, --n-poses 9, and --cpu 1 per clone.
References
Run this workflow
The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.