October 5, 2026 · Temple Compute
Airflow vs Nextflow: Which Fits Scientific Pipelines?
Airflow and Nextflow both draw a DAG and run it. That is roughly where the similarity ends. Airflow is a scheduler for recurring data pipelines. Nextflow is a workflow language for scientific pipelines that run once per dataset, often on a cluster. Teams compare them because a lab already has one of them, usually Airflow from the data team, and wants to know whether it can carry the science too.
Short answer: for compute-heavy scientific work, Nextflow is the better fit of the two. Airflow wins when the job is scheduling, not compute. Detail below.
We build a third option, Temple Compute OS on the open-source Horus engine, and say where it fits at the end. Read that part as a vendor's view.
What each one is for
Apache Airflow orchestrates scheduled work. A DAG runs on a cron schedule or when upstream data lands, and the scheduler tracks every run, retries failures, and backfills missed intervals. Its strength is the provider ecosystem: hundreds of operators for databases, warehouses, queues and SaaS APIs. It was built for ETL, and it does ETL very well.
Nextflow runs data-driven scientific pipelines. You launch a run against a dataset, files flow through processes via channels, and the run ends. It was built for bioinformatics, and nf-core gives it a large, peer-reviewed catalogue of ready pipelines. Seqera adds a commercial platform on top for launching, monitoring and compute environment management.
Pipeline definition: Python DAGs vs a Groovy DSL
Airflow DAGs are plain Python. Tasks are operators or decorated functions, and
dependencies are set in code. Any Python developer can read one. Dynamic task
mapping (.expand()) lets a task fan out over a list produced at run time.
Nextflow pipelines are written in a Groovy-based DSL. Processes declare inputs, outputs and a script block. Channels connect them, and the dataflow model decides what runs when. It is concise once learned, and it is a real language with its own idioms and debugging experience. For many labs "who can edit the Nextflow" becomes a staffing question.
Scheduling model
The two answer different questions.
- Airflow asks when should this run? It is a long-running service with a scheduler, a metadata database, a web UI and workers. Runs are tied to schedule intervals or data events.
- Nextflow asks what can run now? Each run is a process you launch. Tasks start as soon as their input files exist. There is no built-in cron; you trigger runs yourself or through Seqera.
Airflow passes small values between tasks (XCom) and expects big data to live in external storage. Nextflow stages files between tasks as a core feature, in a work directory per task.
HPC and SLURM support
This is the biggest practical gap.
Nextflow has native executors for SLURM, PBS/Torque, PBS Pro, LSF, SGE and
HTCondor, plus Kubernetes and the major cloud batch services. Per-process
cpus, memory, time and queue directives map onto the scheduler. Moving
a pipeline to a cluster is mostly configuration.
Airflow has no official HPC scheduler executor. Its executors (Local,
Celery, Kubernetes) run tasks on workers you operate. Reaching a cluster means
a custom operator, or SSH plus sbatch, plus polling for completion and
handling failures yourself. Many labs have written that code. Few enjoy
maintaining it.
Containers and environments
Nextflow supports Docker, Singularity/Apptainer, Podman and conda per process. Different stages can use different images in one pipeline, which matters when tools have conflicting dependencies.
Airflow runs tasks in the worker's environment by default. Per-task isolation comes from operators such as the Docker or Kubernetes pod operators, or virtualenv-based operators. It works, but it is per-operator wiring rather than a pipeline-wide model.
Where each breaks down for scientific compute
Airflow:
- It does not provide compute. You bring and run the workers.
- No native HPC scheduler support.
- The task model favours short, uniform, idempotent jobs. Multi-day simulations and very wide fan-outs fit badly.
- It tracks task success, not what a task produced.
Nextflow:
- The DSL is a barrier for scientists who don't write Groovy.
- Compute underneath is still yours: cluster config, cloud accounts, queues.
- The ecosystem is overwhelmingly bioinformatics. Materials, CFD or chemistry pipelines get less from nf-core.
- Iterative loops (run, evaluate, run again) are awkward; recursion is still a preview feature.
Comparison table
| Apache Airflow | Nextflow / Seqera | |
|---|---|---|
| Built for | Scheduled data pipelines (ETL) | Reproducible scientific pipelines |
| Pipeline definition | Python | Groovy-based DSL |
| Execution model | Long-running scheduler service | Per-run, dataflow-driven |
| Cron / scheduled runs | Excellent | Not built in (Seqera adds it) |
| HPC schedulers | Custom operators only | Native: SLURM, PBS, LSF, SGE, more |
| Per-task containers | Via specific operators | Native, per process |
| File passing between tasks | External storage; XCom for small values | Native, staged per task |
| Resume after failure | Clear and rerun failed tasks | -resume from cached tasks |
| Fan-out | Dynamic task mapping | Channels, native |
| Provides compute | No | No (Seqera manages environments you own) |
| Ecosystem | Huge integration catalogue | nf-core, bioinformatics |
| Licence | Apache 2.0 | Apache 2.0 |
When to pick which
- Pick Airflow if the work is scheduled data movement: ingest instrument output nightly, load a warehouse, trigger reports. Keep it for that even if the science runs elsewhere.
- Pick Nextflow if the work is bioinformatics, nf-core already has your pipeline, someone on the team owns the DSL, and you have a cluster or batch environment to run on.
- Use both when it fits: Airflow on a schedule triggers a Nextflow run. This is a common, reasonable pattern.
If neither fits, the cause is usually one of three things: nobody can maintain the DSL, the domain isn't bioinformatics, or nobody owns the compute underneath.
Where Temple Compute OS fits
Temple Compute OS targets that third case. Each stage of a workflow declares its own compute target: local, SSH, an HPC scheduler, or cloud. The engine moves artifacts between stages when an edge crosses machines. Workflows are built visually or in YAML, with a Python API for those who want it, so there is no DSL. Completed tasks are skipped on re-run, so a failed run picks up where it stopped. Fan-out over a collection, bounded loops, subworkflows and runtime DAG mutation are built into the engine.
It is younger than both. It has nothing like Airflow's integration catalogue or nf-core's pipeline coverage, and its cron scheduling is basic. The full trade-offs are on the Airflow and Nextflow / Seqera comparison pages, and the wider field is in our orchestration tools roundup.
FAQ
Is Nextflow better than Airflow for HPC job orchestration?
For HPC, yes. Nextflow submits to SLURM, PBS, LSF and SGE natively and maps per-process resources onto the scheduler. Airflow needs a custom operator to reach a cluster.
Can Airflow run bioinformatics or simulation pipelines?
It can, and some teams do. You supply the workers, the containers and the cluster integration yourself, and long-running or very wide stages fight the scheduler's assumptions.
Can I use Airflow and Nextflow together?
Yes. A common pattern is an Airflow DAG that triggers a Nextflow run on a schedule or when new data arrives, then picks up the outputs.
What is an orchestration platform for HPC that needs no DSL?
Temple Compute OS routes each workflow stage to HPC, cloud or local compute, with workflows built visually, in YAML or in Python. The engine, Horus, is open source.
Open Horus to build a workflow across HPC and cloud.