All posts

October 5, 2026 · Temple Compute

How to Automate Parametric Studies on HPC and Cloud

Most parametric studies start as a shell loop. Edit a value, submit a job, repeat. Then a job array, then a merge script, then a spreadsheet tracking which of the 200 runs failed and why. By the time the study is done, a meaningful share of the work went into plumbing.

The plumbing is the same whatever the field: a CFD mesh sweep, a materials screen over compositions, a molecular simulation over replicas or mutations. This post breaks a parametric study into the few shapes an orchestrator has to support, and shows how they map onto a real workflow.

The shape of a parametric study

Almost every study is built from four parts:

  1. Prepare. Generate inputs: a design of experiments, a set of structures, a mesh per geometry. Usually cheap and CPU-bound.
  2. Fan out. Run the same simulation once per parameter set. Each run is independent. This is where the compute goes.
  3. Gather. Collect every run's output into one place and reduce it: a table, a response surface, an estimate with error bars.
  4. Decide. Stop, or pick new parameters and go again. Only optimization studies have this step.

An orchestrator that handles parametric work well makes steps 2 and 3 a declaration rather than code, and keeps step 4 inside the same run.

Sweeps are a fan-out map plus a gather

A sweep over a collection is a map: one stage definition, cloned once per element. In the Horus engine behind Temple Compute OS this is the map: block. A producer stage writes the collection; the map creates one clone per element, each in its own working directory; the clones run concurrently; the outputs land under one <stage>.gathered/<i>/ folder that a gather stage reads.

Two properties matter for parametric work:

  • The clone count follows the data. Add rows to the design of experiments and you get more clones. The graph doesn't change.
  • The gather stage doesn't know the count. It reads one folder. You don't maintain a list of 200 expected output paths.

The minimal version is the fan-out, map and gather showcase: split a list into batches, map a stage over the batches, gather the results.

Replicas and seeds are a bounded loop

Sometimes there is no collection, just a count: 20 replicas, 50 random seeds, 10 repeats of a stochastic method. Faking a collection with a producer that emits 50 placeholder files is pure overhead.

Horus handles this with map: and range: N. The engine creates clones 0 to N-1 with no producer stage, and each clone reads its own index to pick a seed or a parameter value. The loop is bounded, so the engine can plan the whole run in advance. See the bounded loop showcase.

Optimization loops: a round at a time

Optimization adds the decide step: evaluate a batch, choose the next batch from the results, repeat until it converges or the budget runs out. The next round's parameters don't exist until the current round finishes.

Be precise about what a bounded range: loop is: N clones that run concurrently, not N rounds that feed each other. For round-based optimization, each round is a map, and the decision between rounds is a stage. Horus has a runtime DAG-mutation API for this: a stage can read its inputs and add new tasks and edges to the running workflow from inside its own function. A decision stage can read round k's gathered results and append round k+1, or append nothing and let the run end. The programmatic dynamic DAG showcase shows the pattern on toy data.

The optimizer itself, whether Bayesian, gradient-free or a simple grid refinement, is your code in that stage. The engine provides the loop structure, not the algorithm.

Mixed targets: CPU, GPU and HPC in one study

The stages of a study rarely want the same hardware. Input preparation runs in seconds on a CPU. The simulations want GPU nodes or a large share of an HPC allocation. The gather and analysis are a Python script.

Running the whole study on one machine type wastes money or time. On the GPU node you pay GPU rates to parse text files. On a workstation the wide band runs in series.

In Horus each stage declares its own executor and target, and the engine moves artifacts between them when an edge crosses machines. The prep stage stays local, the mapped simulation stage goes to an HPC scheduler or cloud, the gather comes back. Moving a stage is a change to its executor: or target: block, not to its command.

Skip the runs that already finished

A 200-run study will lose some runs: a node fails, a wall-time limit hits, one parameter set diverges. Rerunning the whole sweep for a handful of failures is the most expensive mistake in parametric work.

Horus skips completed tasks on re-run. A task whose output artifacts already exist is not executed again, so a study that failed at run 180 resumes at run 180. Mapped clones are separate tasks with their own outputs, so the same check applies to each one. The same mechanism means that editing only the analysis stage reruns only the analysis. To force a fresh run, delete the outputs or pass --no-skip for a task or --no-skip-all for everything.

Worked example: mutation free energy

The mutation free energy workflow in our open-source library has exactly this shape. It computes a ΔΔG for an Ile10 to Ala mutation in staphylococcal nuclease with pmx and GROMACS, using fast-growth non-equilibrium thermodynamic integration.

  • Prepare (CPU, seconds). Extract snapshots from the wild-type and mutant trajectories, model the hybrid structures, build the dual topology.
  • Simulate (GPU or HPC, the wide band). Minimize, equilibrate, then run TI transitions in both directions. One transition is a short GROMACS run. The estimators need dozens per direction to converge, and a real study runs 80 or more. They share nothing, so they are a fan-out.
  • Gather (CPU, seconds). pmxanalyse reads the accumulated work values and estimates the free energy with the Crooks Gaussian Intersection, BAR and Jarzynski estimators.

The library version runs one transition per direction to show the method and ships pre-computed replicate work values for the analysis. Scaling it to a production sample is the fan-out pattern above: map the TI stage over the snapshot collection, send that stage to a GPU or HPC executor, and keep prep and analysis local. Change the estimator options and only pmxanalyse reruns.

The same structure fits MD setup sweeps. The GROMACS MD setup and protein-ligand complex MD setup workflows are natural map bodies when the parameter is the input structure, the ligand or the force-field settings.

What to look for in an orchestrator

For parametric studies on HPC or cloud, check for:

  • Declarative fan-out over a collection, with fan-out width set by the data.
  • A count-based loop for replicas and seeds.
  • A way to add rounds at run time, for optimization.
  • Per-stage compute targets, including your institution's scheduler.
  • Restart that skips completed runs, at the granularity of one run.
  • Data movement between stages handled by the engine, not by you.

FAQ

How do I automate parametric studies and optimization pipelines on cloud HPC?

Model the study as prepare, fan out, gather and decide. Use an orchestrator that maps one stage over the parameter sets, gathers the outputs, routes the simulation stage to HPC or cloud, and skips finished runs on restart.

What is the best way to orchestrate CAE and simulation workflows across HPC and cloud?

Use per-stage compute targets. Keep cheap pre- and post-processing on CPU, send the solver stage to an HPC scheduler or cloud GPUs, and let the orchestrator move the files between them.

How is a parameter sweep different from an optimization loop?

A sweep knows all its parameter sets up front, so every run starts at once. An optimization loop picks the next parameters from the last results, so it runs in rounds, with a decision step between them.

Do I need to rerun the whole sweep when a few jobs fail?

Not if the orchestrator skips completed work. In Horus, a task whose outputs already exist is skipped on re-run, so only the failed or missing runs execute.

Open Horus to run your next sweep across HPC and cloud.