Molecular dynamics
fpocket Cavity Analysis Across Conformational Ensembles
Run fpocket on every structure in a conformational ensemble, filter cavities by score, druggability and volume, and rank models by their best cavity.
What this workflow does
This workflow detects and ranks druggable cavities across a conformational ensemble. The input is a folder of related structures: MD cluster representatives, NMR models or any set of conformations. The bundled example holds three MD cluster representatives.
For each model, the workflow extracts the protein and runs fpocket. fpocket_filter keeps the cavities whose score, druggability and volume fall inside the configured windows. An optional residue filter keeps only the cavities whose centre of mass sits near a residue selection, for a known site.
A final step ranks the models by their best cavity. It writes three orderings: by volume, by druggability score and by fpocket score. The metrics often disagree, and which one matters depends on what the cavity is for.
On the bundled example, 3 models are analysed, 2 keep a cavity and 4 cavities
survive in total. cluster1 wins on all three metrics. cluster2 has no
druggable cavity: its best druggability score is 0.033.
A protein is not one shape. A pocket that is closed in the crystal structure can be open in part of the ensemble. This workflow tells you which conformation to dock into. Feed the winner to cavity-guided virtual screening.
This is a port of the cavity_analysis CLI from biobb_vs_workflows.
The compute problem
The models are independent. The source workflow still analyses them in a serial loop. Each model takes about 30 seconds, so a 100-model ensemble is 100 independent tasks: a natural cluster array job.
A single failing model also aborts the source run. With a large ensemble, one bad structure costs every result.
How Horus solves it
The analysis stage is a map: block. Horus starts one concurrent clone per
model, and each clone runs extraction, fpocket, the metric filter and the residue
filter. A gather step collects every clone's output for ranking. Scaling is
linear in ensemble size and lives entirely in the map stage.
To move the analysis onto a cluster, give the map template a target: and a
resources: block. Nothing else changes.
A model with no surviving cavity is a result, not a failure. A model that errors
is recorded in cavity_report.json and the run continues.
Every step calls the BioBB Python API in-process, so fpocket comes from the conda environment. No container is needed.
Pipeline
prepare_models Wrap each PDB in its own model_NNN/ folder
│
analyze[00..NN] One concurrent clone per model:
│ extract_molecule Keep the protein (optionally chains)
│ fpocket_run Detect cavities
│ fpocket_filter Filter by score, druggability, volume
│ filter_residue_com Keep cavities near a residue selection
│
summarize Rank models by their best cavity
Inputs and outputs
Inputs
examples/structures/: a folder of PDB models. A single.pdbfile also works.
Outputs land in horus_workflow_results/results/:
summary_by_volume.yml: models ordered by their largest cavity.summary_by_drug_score.yml: models ordered by their most druggable cavity.summary_by_score.yml: models ordered by their best-scoring cavity.cavity_report.json: models analysed, models with a cavity, the winner on each metric and any failures.
Per-model detail stays in analyze.gathered/<i>/: every cavity fpocket found,
the survivors and the prepared structure.
The filter windows live in the analyze template: --score, --druggability
and --volume. The bundled windows (0,1, 0.2,1, 100,5000) are looser than
the source defaults, which reject every cavity in this ensemble. Tighten them
for a larger or better-formed set. Add --residue-selection and
--distance-threshold to target a known site, or --chains to keep only some
chains.
The workflow starts from an already clustered ensemble. To build one, see protein conformational ensembles.
Run the workflow
Install the horus-runtime one time:
uv sync
If you do not have uv, install it first:
curl -LsSf https://astral.sh/uv/install.sh | sh
You can also install the packages with pip:
pip install horus-runtime horus-environments
Then run the workflow:
uv run horus run workflow.yaml
The first run builds the conda environment. This takes a few minutes. Later runs
reuse the cached environment at ../.horus_conda_env_cavity. On Apple Silicon,
prefix the run with CONDA_SUBDIR=osx-64 CONDA_OVERRIDE_OSX=10.16.
References
Run this workflow
The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.