October 5, 2026 · Temple Compute
fpocket Cavity Analysis for Cavity-Guided Virtual Screening
Structure-based virtual screening needs a docking box. If the binding site is annotated, you draw the box around it. If it is not, you have to find the pocket first, and the structure you find it on decides what you dock into.
This post walks through two open workflows from our workflow library that handle both steps with fpocket and AutoDock Vina, built on BioExcel Building Blocks:
- fpocket cavity analysis across a conformational ensemble detects and ranks cavities on every model and tells you which conformation to use.
- Cavity-guided virtual screening detects cavities on a receptor, picks the most druggable one, builds the docking box around it and screens a ligand library.
Both are ports of the cavity_analysis and vs_autodock CLIs from
biobb_vs_workflows.
Step 1: detect and rank cavities across an ensemble
A protein is not one shape. A pocket that is closed in the crystal structure can be open in part of the ensemble. So binding pocket detection with fpocket should run on more than one conformation.
The cavity analysis workflow takes a folder of related structures: MD cluster representatives, NMR models, or any set of conformations. For each model it:
- Extracts the protein, optionally only some chains.
- Runs fpocket to detect cavities.
- Filters cavities by fpocket score, druggability score and volume.
- Optionally keeps only cavities whose centre of mass sits within a distance threshold of a residue selection, for a known site.
A final step ranks the models by their best surviving cavity and writes three orderings: by volume, by druggability and by score. Three orderings exist because the metrics disagree. A large cavity is not always a druggable one.
On the bundled example of three MD cluster representatives, two models keep a
cavity and four cavities survive in total. cluster1 wins on all three
metrics. cluster2 is reported as having no druggable cavity: its best
druggability score is 0.033. That is a result, not a failure, and it is the
kind of answer you want before you spend compute docking into the wrong shape.
If you do not have an ensemble yet, the protein conformational ensembles workflow generates one and clusters it into representatives.
The filter windows matter more than they look. The source defaults (score and druggability at least 0.4, volume at least 200 ų) reject every cavity in the bundled ensemble. The workflow ships looser windows and says so. Tighten them for a larger or better-formed set.
Step 2: use the top cavity to define the docking box
The cavity-guided virtual screening workflow runs the screen. The example target is p38-alpha MAP kinase (PDB 3HEC) with a bundled library of 8 ZINC compounds.
The pipeline:
- Fetch the receptor and keep only the protein.
- Run fpocket and filter the cavities by score, druggability and volume.
- Rank the survivors by druggability score, fpocket score or volume, with an optional residue-proximity filter. The winner is written as the config for the next step, so the pocket number is never retyped by hand.
- Extract the winning pocket and build a docking box around it.
- Protonate the receptor and convert it to PDBQT.
- Split the library into batches and dock each batch with AutoDock Vina.
- Rank every ligand by its best affinity and write the top poses.
On the bundled example, the selection step picks pocket 1 with a druggability score of 0.876. All 8 ligands dock. The best affinity is -5.157 kcal/mol.
The same receptor shows why one metric is not enough. Its ATP-site cavity
scores 0.341 on fpocket's score despite that 0.876 druggability, so a strict
score window would throw away the real site. The workflow ranks by
druggability by default and lets you switch with --rank-by.
The box is kept tight on purpose: a 5 Å offset between the last cavity atom and the box boundary. Vina cannot place a ligand outside the box.
For one ligand instead of a library, the protein-ligand docking with fpocket workflow runs the same chemistry end to end and writes the final complex.
The limits of geometric pocket detection
fpocket is geometric. It finds cavities by filling the surface of the structure you give it with alpha spheres, then scores the clusters. That has consequences.
- It only sees the shape in front of it. A pocket that is closed in the input structure is invisible. Running across an ensemble helps, but only for conformations the ensemble actually samples.
- Cryptic pockets need more than this. Some sites only open with dynamics or in the presence of a ligand. A static ensemble of cluster representatives may miss them. We cover that problem in cryptic pocket detection, the follow-up to this post.
- Scores are heuristics. Volume, score and druggability can disagree on
the same pocket, as 3HEC shows. Pick the metric that matches what the pocket
is for, and check the full
pocket_report.json, not only the winner. - A good pocket is not a good hit. The docking score ranks ligands inside the box you chose. If the box sits on the wrong cavity, every ranking that follows is wrong too.
When you know the site, say so. Both workflows take a residue selection and a distance threshold, so detection stays automatic but cannot wander to a surface groove on the other side of the protein.
How Horus runs it
Both workflows run on Horus, the open-source engine behind Temple Compute OS.
The independent work fans out. In the cavity analysis workflow, every
model is a concurrent clone. The source runs them in a serial loop. At about
30 seconds per model, a 100-model ensemble becomes 100 independent tasks. In
the screening workflow, the source docks ligands one after the other. Here the
library is split into batches and each batch docks in its own clone.
--batch-size is the dial: one ligand per clone gives the most parallelism,
at the cost of an environment start-up per ligand.
Each stage picks its own compute. Only docking scales with library size.
Fetching, cavity detection and pocket selection run once and take seconds. The
8-ligand example runs in about two minutes on a laptop. A 100k-compound
library is a cluster job, and moving it there means giving the docking stage a
target: and resources:. Nothing else in the pipeline changes.
Install problems stay contained. In the screening workflow, the
fpocket-backed stages run inside the biobb_vs Linux container and the rest
runs in a native conda environment.
One bad input does not cost the run. A ligand that fails to convert or
dock is recorded in screening_summary.json with its stage and error. A model
that fails cavity analysis is recorded in cavity_report.json. The rest of
the screen keeps going.
Today the two workflows are chained by hand: you read the winning model from the cavity analysis and point the screen at it.
FAQ
How does fpocket rank binding pockets?
fpocket reports a score, a druggability score and a volume for every cavity it detects. Both workflows let you filter by windows on all three and rank the survivors by one of them. The cavity analysis workflow writes one summary per metric, because they often disagree.
Can fpocket define the docking box for AutoDock Vina?
Yes. The cavity-guided screening workflow extracts the selected fpocket pocket and builds a box around it with a configurable offset, 5 Å by default. Vina then docks every ligand inside that box.
Should I run fpocket on one structure or an ensemble?
An ensemble, if you have one. Pockets open and close between conformations. Running cavity analysis across cluster representatives or NMR models tells you which conformation has the best cavity before you screen.
Does fpocket find cryptic pockets?
Only if the pocket is open in a structure you give it. fpocket is geometric and works on static structures. Sites that open only with dynamics or ligand binding need a different approach, covered in cryptic pocket detection.
Run it
Both workflows are open source in the workflow library. To run them with per-stage compute and no scheduler config, open Horus.