All workflows

Molecular dynamics

fpocket Cavity Analysis Across Conformational Ensembles

Run fpocket on every structure in a conformational ensemble, filter cavities by score, druggability and volume, and rank models by their best cavity.

fpocketMDAnalysisBioExcel biobbbiobb_vsbiobb_structure_utils

What this workflow does

This workflow detects and ranks druggable cavities across a conformational ensemble. The input is a folder of related structures: MD cluster representatives, NMR models or any set of conformations. The bundled example holds three MD cluster representatives.

For each model, the workflow extracts the protein and runs fpocket. fpocket_filter keeps the cavities whose score, druggability and volume fall inside the configured windows. An optional residue filter keeps only the cavities whose centre of mass sits near a residue selection, for a known site.

A final step ranks the models by their best cavity. It writes three orderings: by volume, by druggability score and by fpocket score. The metrics often disagree, and which one matters depends on what the cavity is for.

On the bundled example, 3 models are analysed, 2 keep a cavity and 4 cavities survive in total. cluster1 wins on all three metrics. cluster2 has no druggable cavity: its best druggability score is 0.033.

A protein is not one shape. A pocket that is closed in the crystal structure can be open in part of the ensemble. This workflow tells you which conformation to dock into. Feed the winner to cavity-guided virtual screening.

This is a port of the cavity_analysis CLI from biobb_vs_workflows.

The compute problem

The models are independent. The source workflow still analyses them in a serial loop. Each model takes about 30 seconds, so a 100-model ensemble is 100 independent tasks: a natural cluster array job.

A single failing model also aborts the source run. With a large ensemble, one bad structure costs every result.

How Horus solves it

The analysis stage is a map: block. Horus starts one concurrent clone per model, and each clone runs extraction, fpocket, the metric filter and the residue filter. A gather step collects every clone's output for ranking. Scaling is linear in ensemble size and lives entirely in the map stage.

To move the analysis onto a cluster, give the map template a target: and a resources: block. Nothing else changes.

A model with no surviving cavity is a result, not a failure. A model that errors is recorded in cavity_report.json and the run continues.

Every step calls the BioBB Python API in-process, so fpocket comes from the conda environment. No container is needed.

Pipeline

prepare_models               Wrap each PDB in its own model_NNN/ folder
   │
analyze[00..NN]              One concurrent clone per model:
   │                           extract_molecule    Keep the protein (optionally chains)
   │                           fpocket_run         Detect cavities
   │                           fpocket_filter      Filter by score, druggability, volume
   │                           filter_residue_com  Keep cavities near a residue selection
   │
summarize                     Rank models by their best cavity

Inputs and outputs

Inputs

  • examples/structures/: a folder of PDB models. A single .pdb file also works.

Outputs land in horus_workflow_results/results/:

  • summary_by_volume.yml: models ordered by their largest cavity.
  • summary_by_drug_score.yml: models ordered by their most druggable cavity.
  • summary_by_score.yml: models ordered by their best-scoring cavity.
  • cavity_report.json: models analysed, models with a cavity, the winner on each metric and any failures.

Per-model detail stays in analyze.gathered/<i>/: every cavity fpocket found, the survivors and the prepared structure.

The filter windows live in the analyze template: --score, --druggability and --volume. The bundled windows (0,1, 0.2,1, 100,5000) are looser than the source defaults, which reject every cavity in this ensemble. Tighten them for a larger or better-formed set. Add --residue-selection and --distance-threshold to target a known site, or --chains to keep only some chains.

The workflow starts from an already clustered ensemble. To build one, see protein conformational ensembles.

Run the workflow

Install the horus-runtime one time:

uv sync

If you do not have uv, install it first:

curl -LsSf https://astral.sh/uv/install.sh | sh

You can also install the packages with pip:

pip install horus-runtime horus-environments

Then run the workflow:

uv run horus run workflow.yaml

The first run builds the conda environment. This takes a few minutes. Later runs reuse the cached environment at ../.horus_conda_env_cavity. On Apple Silicon, prefix the run with CONDA_SUBDIR=osx-64 CONDA_OVERRIDE_OSX=10.16.

References

Run this workflow

The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.