All workflows

AI and ML

Train an LLM From Scratch: Pretrain, SFT and Eval

Train a small LLM from scratch in one DAG. Tokenize Pile text, pretrain, fine-tune on Alpaca, Dolly and GSM8K, then score GSM8K. Reruns skip done stages.

PyTorchtiktokendatasetsh5pyconda

What this workflow does

This workflow trains a small decoder-only Transformer from scratch. It clones the train-llm-from-scratch repo, which is pure PyTorch with no transformers, trl, or peft. It then runs the full chain: tokenize, pretrain, supervised fine-tune (SFT), and evaluate.

The pretraining data is a shard of the Pile. The SFT data is Alpaca, Dolly, and GSM8K, rendered with a chat template and a loss mask, then packed. The eval stage greedily decodes the SFT checkpoint on held-out GSM8K questions and records its accuracy.

Two more data stages build inputs for later work. One builds preference pairs from HH-RLHF and UltraFeedback for a reward model or DPO. The other builds GSM8K and arithmetic prompt sets for PPO or GRPO. No task in this workflow consumes them yet.

The workflow is a work in progress. Treat it as a reference pipeline.

The compute problem

An LLM training run has stages with very different hardware needs.

The clone stage needs git and a few seconds.

The data stages stream datasets and tokenize them into HDF5. They are CPU and network work. They run in parallel with each other, and they need no GPU.

The pretrain and SFT stages are the expensive ones. At the smoke-scale config they run in minutes on a laptop. At the repo's default ~400M-parameter config they need a real GPU and hours to days.

The eval stage loads one checkpoint and decodes. It needs far less than training.

Most setups run all of this as one long shell script on one machine. You hold a GPU while tokenization runs. A crash in SFT means you rerun the data prep, or you comment out lines by hand to skip it.

How Horus solves it

Horus runs each stage as its own task with declared inputs and outputs. The edges: wire one task's output file to the next task's input. A task starts only once its upstream artifacts exist.

Each task has its own target:. The shipped YAML runs every task locally. To send training to a GPU machine, change the target: of train_pretrain and train_sft. The data stages stay on CPU. The runtime.command string does not change, and Horus moves the HDF5 files and checkpoints across the boundary.

Each stage is resumable. On a rerun, a stage whose outputs already exist is skipped. Change the SFT config and only SFT and eval run again. The tokenized data and the base checkpoint are reused.

Horus builds the Python environment from conda_env.yaml with micromamba. You do not install PyTorch by hand on each host.

Pipeline

clone_repo (local, shell)              git clone            ──► repo/
   │
   ├─► prepare_pretrain_val   (CPU)    Pile val shard       ──► pile_dev.h5
   ├─► prepare_pretrain_train (CPU)    Pile train shards    ──► pile_train.h5
   │        └─► train_pretrain (GPU)   pretrain base model  ──► base_pretrained.pt
   │                 └─► train_sft (GPU)  SFT               ──► sft.pt
   │                          └─► eval_sft  GSM8K accuracy  ──► stage_table.jsonl
   │
   ├─► prepare_sft_data        (CPU)   Alpaca/Dolly/GSM8K   ──► sft_packed.h5
   ├─► prepare_preference_data (CPU)   HH-RLHF/UltraFeedback ─► preferences.jsonl
   └─► prepare_rl_prompts      (CPU)   GSM8K + arithmetic   ──► rl_prompts_*.jsonl

Inputs and outputs

Inputs

None. The workflow clones its own training repo. It downloads the Pile shard, Alpaca, Dolly, GSM8K, HH-RLHF, and UltraFeedback from Hugging Face.

Outputs land in workflow_results/:

  • ckpts/base_pretrained.pt: the pretrained base model.
  • ckpts/sft.pt: the fine-tuned model.
  • logs/pretrain.jsonl, logs/sft.jsonl: one JSON record per logged training step.
  • logs/stage_table.jsonl: the GSM8K accuracy row for the SFT checkpoint.
  • data/: the tokenized datasets, plus the preference and RL prompt sets.

Each checkpoint embeds its model config, training step, and metrics. You need no separate architecture file to load it.

Run the workflow

Install the horus-runtime and the plugins one time:

uv sync

If you do not have uv, install it first:

curl -LsSf https://astral.sh/uv/install.sh | sh

You can also install the packages with pip:

pip install horus-runtime horus-environments

Then run the workflow:

uv run horus run workflow.yaml

To set the model scale, edit --config in the train_pretrain and train_sft commands. The default is configs/smoke/{pretrain,sft}.json: a tiny model (n_embed=64, n_head=4, n_blocks=2) trained for about 10 to 20 steps. Point it at configs/pretrain.json and configs/sft.json for the repo's default ~400M-parameter model.

To override a training field, add --field value to the task's command:. Any field on PretrainConfig or SFTConfig works, for example --lr, --batch_size, --train_steps, or --use_wandb true --wandb_project <name>.

To pull more pretraining text, raise --num_shards on prepare_pretrain_train. The default is 1.

For a faster first run on a smaller dataset, start with the TinyStories workflow.

References

Run this workflow

The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.