All workflows

Molecular dynamics

fpocket + AutoDock Vina Cavity-Guided Virtual Screening

Detect cavities with fpocket, pick the most druggable one and dock a ligand library into it with AutoDock Vina. Docking runs as parallel batches.

fpocketAutoDock VinaOpenBabelMDAnalysisBioExcel biobbbiobb_vs

What this workflow does

This workflow runs an end-to-end structure-based virtual screen. It fetches a receptor, finds its cavities with fpocket, picks one pocket and docks a ligand library into that pocket with AutoDock Vina. The ligands are ranked by binding affinity. The target is p38-alpha MAP kinase (PDB 3HEC). The bundled library holds 8 ZINC compounds.

The workflow strips the receptor down to the protein and runs fpocket on it. fpocket_filter keeps the cavities whose score, druggability and volume fall inside the configured windows. A selection step ranks the survivors by druggability score, fpocket score or volume and names the winner. You can also keep only the cavities near a residue selection, for a known site. The workflow builds a docking box around the winning pocket.

In parallel, the receptor gets hydrogens and becomes a PDBQT file. The library is split into batches. Each batch is docked with AutoDock Vina, and a final step ranks every ligand by its best affinity and writes the top poses.

On the bundled example, fpocket selects pocket 1 (druggability 0.876). All 8 ligands dock. The best affinity is -5.157 kcal/mol.

This is a port of biobb_vs_workflows, which ships the same science as two separate CLIs. Here they are one pipeline. The pocket number flows from detection into docking and is never retyped by hand.

The compute problem

Docking is embarrassingly parallel and linear in library size. The source workflow docks ligands one after the other, with no parallelism between them. The 8-ligand example runs in about two minutes on a laptop. A 100k-compound library is a cluster job.

Only the docking stage scales. Receptor fetch, cavity detection and pocket selection run once on the receptor and take seconds. The two halves want different machines.

The tools also have awkward install profiles. biobb_vs pins fpocket 4.1, which has no osx-arm64 conda build. On Apple Silicon the environment must build under Rosetta.

How Horus solves it

Horus runs the fpocket-backed stages inside the biobb_vs Linux container. The rest runs in one native conda environment. Only the executor: field changes between them.

The docking stage is a map: block. The library splits into batch_NNN/ folders and Horus starts one concurrent clone per batch. Each batch carries its own copy of the prepared receptor and the docking box, so a clone is self-contained and can run on any target. To move docking onto a cluster, give the map template a target: and a resources: block. Nothing else in the pipeline changes.

--batch-size is the dial. One ligand per clone gives the most parallelism and the finest resume granularity, but pays an environment start-up per ligand.

A ligand that fails to convert or dock is recorded, not fatal. The docking script never exits non-zero, so one bad molecule cannot cost the whole screen. Failures land in screening_summary.json with the stage and the error.

Pipeline

fetch_pdb                    Download 3HEC structure from PDB
   │
extract_protein               Keep the protein (drop water, ions, ligands)
   │
   ├─ fpocket_run              Detect cavities with fpocket
   │     │
   │  fpocket_filter           Filter cavities by score, druggability, volume
   │     │
   │  select_pocket            Rank survivors, write the fpocket_select config
   │     │
   │  fpocket_select           Extract the winning pocket
   │     │
   │  box                      Build the docking box around it (offset 5 Å)
   │
   └─ str_check_add_hydrogens  Protonate receptor, convert to PDBQT
   │
split_ligands                 Split the library into batch_NNN/ folders
   │
dock[00..NN]                  AutoDock Vina, one concurrent clone per batch
   │
rank                          Rank ligands by affinity, write top poses

Inputs and outputs

Inputs

  • configs/fetch_pdb.yaml: the PDB code to screen against (default 3HEC).
  • examples/ligands.sdf: the ligand library. A .smi file also works.
  • configs/*.yaml: one config per BioBB step.

Outputs land in horus_workflow_results/results/:

  • scores.csv: the ranking, most negative affinity first.
  • screening_summary.json: docked and failed counts, the best hit and every failure.
  • poses/<ligand>_poses.pdb: the top poses as PDB.
  • pocket_report.json: every candidate cavity, which survived filtering and why the winner won.

To change the pocket metric, set --rank-by on select_pocket. To restrict to a known site, add --residue-selection and --distance-threshold. The bundled filter windows are loose on purpose: the 3HEC ATP-site cavity scores 0.341 on fpocket's score despite its 0.876 druggability. Tighten configs/fpocket_filter.yaml for a holo structure.

Run the workflow

Install the horus-runtime one time:

uv sync

If you do not have uv, install it first:

curl -LsSf https://astral.sh/uv/install.sh | sh

You can also install the packages with pip:

pip install horus-runtime horus-environments

Docker is required for the fpocket stages. Then run the workflow:

uv run horus run workflow.yaml

The first run builds the conda environment. Later runs reuse the cached environment at ../.horus_conda_env_vs. On Apple Silicon, prefix the run with CONDA_SUBDIR=osx-64 CONDA_OVERRIDE_OSX=10.16.

To pick the receptor conformation first, run fpocket cavity analysis across an ensemble. For a single ligand, see protein-ligand docking with fpocket.

References

Run this workflow

The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.