All workflows

Drug discovery

Lacuna + AutoDock Vina: Cryptic Pocket Docking

Find cryptic pockets with Lacuna, then dock a ligand library into a ranked site with AutoDock Vina. Meeko and OpenBabel prep, one Vina run per ligand.

LacunaAutoDock VinaMeekoOpenBabelRDKituvconda

What this workflow does

This workflow takes one unbound (apo) receptor structure and a ligand library. It finds the receptor's pockets with Lacuna, including cryptic pockets that are closed or shallow in the input. It then docks every ligand into a ranked pocket with AutoDock Vina. It returns a ranked pocket report and ranked binding energy tables in CSV format.

Lacuna builds a conformational ensemble from the input structure. It detects pockets in every conformer and clusters them across the ensemble. A site that opens in only a few conformers is still reported. A fitted model ranks the clusters and scores how much each site opens relative to the input (its crypticity).

The workflow runs Lacuna with --detector surface-fusion. This pools the geometric detector with a learned surface detector. The geometric detector only proposes a pocket where the surface is already concave. A cryptic site violates that assumption. On held-out CryptoBench data, pooling both detectors takes coverage from 68.5% to 86.4% and top-five recovery from 57.1% to 73.9%.

The prep stage uses OpenBabel for the receptor and Meeko for the ligands. RDKit embeds each SMILES to 3D with ETKDG and MMFF first. AutoDock Vina then docks each ligand into the search box that Lacuna exported for the chosen pocket.

For pocket discovery without docking, see Lacuna Cryptic Pocket Discovery. For docking into a box you define by hand, see AutoDock Vina Docking.

The compute problem

The pipeline has two halves with different needs.

The discovery half generates the ensemble and runs detection on every conformer. This is the expensive part of discovery. It needs only lacuna-pockets, a plain pip package. The default normal-mode backend runs on a CPU.

The docking half needs OpenBabel, Meeko, RDKit, and Vina. These come from conda-forge. The vina package on PyPI has no macOS arm64 wheel. The conda-forge build ships Python bindings for osx-arm64, osx-64, linux-64, and linux-aarch64. One environment for both halves means one dependency solver for two unrelated stacks.

The dock stage scales with the library size. A 100-ligand library is 100 independent Vina runs. Run them one after another and the wall-clock time grows with every ligand you add.

Ensemble generation is also the stage you want to rerun least. If changing the number of exported pockets meant regenerating the ensemble, every small change would cost the full discovery time.

How Horus solves it

Horus assigns an executor to each stage. The discover and dock_prep stages run in a uv-managed venv with lacuna-pockets. The prep and dock stages run in conda-forge environments with the docking stack. The two stacks never share a solver.

The dock stage is a horus_map task. It runs one Vina docking per ligand PDBQT, concurrently. Docking time scales with concurrency, not with the ligand count. Every clone gets a copy of the shared receptor and box.

Discovery and docking-input export are split into two tasks. You change --top or --format on dock_prep without rerunning ensemble generation and detection.

To move a stage to a cluster, give it a target: and resources: block. Give discover one to scale ensemble generation. Give dock one and every clone inherits it. The runtime.command strings do not change. Horus moves the data across each boundary for you.

A ligand that fails to dock does not stop the run. Its clone records status: failed in status.json. The summary stage skips it instead of scoring it as zero.

Pipeline

discover (uv, CPU)     receptor.pdb ──► ensemble + detection + ranking ──► pockets/
   │  pocket_report.json, pocket_*.pdb
dock_prep (uv, CPU)    top 5 pockets ──► docking_inputs/
   │  pocket_*_vina.conf, pocket_*_constraint.yaml, pocket_*_site.pdb
prep (conda, CPU)      receptor + ligands.smi + pocket_0_vina.conf
   │  OpenBabel → receptor.pdbqt, box.json; Meeko → ligands_pdbqt/*.pdbqt
dock[00..NN] (conda)   one AutoDock Vina run per ligand ──► docked/<i>/
   │  {ligand}_out.pdbqt, status.json
summary (CPU)          parse affinities ──► scores.csv + poses.csv
report (CPU)           pocket_report.json ──► report.html + pockets.csv

Inputs and outputs

Inputs

  • receptor.pdb: the unbound target structure. Any single PDB or mmCIF file works. The bundled example is PDB 4LDJ, GDP-bound apo KRAS(G12C) with no switch-II inhibitor present.
  • ligands.smi: one SMILES [name] per line, or a 3D .sdf file. The bundled set is three demo compounds (imatinib, caffeine, aspirin). It proves the docking mechanics. It is not a curated KRAS library.

Outputs land in results/:

  • pockets/pocket_report.json: every surviving pocket cluster with rank, druggability, persistence, crypticity, and contact residues.
  • report.html: a ranked pocket table with a Mol* view of the receptor. pockets.csv holds the same ranking as a flat table.
  • docking_inputs/: Vina box configs, Boltz-2 YAML constraints, and pocket PDBs for the top 5 pockets.
  • receptor_box/ and ligands_pdbqt/: the receptor PDBQT, the resolved box, and one PDBQT per ligand.
  • docked/<i>/: the poses and a status.json for each ligand.
  • scores.csv: one row per ligand, ranked best first: rank, ligand, best_affinity_kcal_mol, mean_affinity_kcal_mol, num_poses.
  • poses.csv: one row per pose: ligand, pose, affinity_kcal_mol, rmsd_lb, rmsd_ub.

A lower affinity value means stronger predicted binding.

On 4LDJ, Lacuna ranks the GDP pocket first. It surfaces the switch-II cryptic pocket twice in the top ten. KRAS is not in the data Lacuna's ranker was fitted on.

Run the workflow

The prep and dock stages build conda environments from conda-forge. A conda family tool must be on your PATH. The workflow.yaml file defaults to micromamba. Set it to mamba or conda if that is what you have.

Install the horus-runtime and the plugins one time:

uv sync

If you do not have uv, install it first:

curl -LsSf https://astral.sh/uv/install.sh | sh

You can also install the packages with pip:

pip install horus-runtime horus-environments

Then run the workflow:

uv run horus run workflow.yaml

To pick the pocket, edit --box-config in the prep stage. The default is pocket_0_vina.conf, Lacuna's rank-1 site. On 4LDJ that is the GDP pocket. Swap in pocket_3_vina.conf to dock into the switch-II cryptic pocket instead. The prep stage keeps only ATOM records, so the bound GDP and waters are removed before docking.

To change discovery, edit the discover stage. Set --conformers for the ensemble size and --rank-by for the ranking strategy. Use --min-druggability, --min-persistence, or --min-crypticity to filter pockets.

To set the search effort, edit the args: of the dock stage. The defaults are --exhaustiveness 8, --n-poses 9, and --cpu 1 per clone.

References

Run this workflow

The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.