AI and ML
Train an LLM From Scratch: Pretrain, SFT and Eval
Train a small LLM from scratch in one DAG. Tokenize Pile text, pretrain, fine-tune on Alpaca, Dolly and GSM8K, then score GSM8K. Reruns skip done stages.
What this workflow does
This workflow trains a small decoder-only Transformer from scratch. It clones the
train-llm-from-scratch
repo, which is pure PyTorch with no transformers, trl, or peft. It then
runs the full chain: tokenize, pretrain, supervised fine-tune (SFT), and
evaluate.
The pretraining data is a shard of the Pile. The SFT data is Alpaca, Dolly, and GSM8K, rendered with a chat template and a loss mask, then packed. The eval stage greedily decodes the SFT checkpoint on held-out GSM8K questions and records its accuracy.
Two more data stages build inputs for later work. One builds preference pairs from HH-RLHF and UltraFeedback for a reward model or DPO. The other builds GSM8K and arithmetic prompt sets for PPO or GRPO. No task in this workflow consumes them yet.
The workflow is a work in progress. Treat it as a reference pipeline.
The compute problem
An LLM training run has stages with very different hardware needs.
The clone stage needs git and a few seconds.
The data stages stream datasets and tokenize them into HDF5. They are CPU and network work. They run in parallel with each other, and they need no GPU.
The pretrain and SFT stages are the expensive ones. At the smoke-scale config they run in minutes on a laptop. At the repo's default ~400M-parameter config they need a real GPU and hours to days.
The eval stage loads one checkpoint and decodes. It needs far less than training.
Most setups run all of this as one long shell script on one machine. You hold a GPU while tokenization runs. A crash in SFT means you rerun the data prep, or you comment out lines by hand to skip it.
How Horus solves it
Horus runs each stage as its own task with declared inputs and outputs. The
edges: wire one task's output file to the next task's input. A task starts
only once its upstream artifacts exist.
Each task has its own target:. The shipped YAML runs every task locally. To
send training to a GPU machine, change the target: of train_pretrain and
train_sft. The data stages stay on CPU. The runtime.command string does not
change, and Horus moves the HDF5 files and checkpoints across the boundary.
Each stage is resumable. On a rerun, a stage whose outputs already exist is skipped. Change the SFT config and only SFT and eval run again. The tokenized data and the base checkpoint are reused.
Horus builds the Python environment from conda_env.yaml with micromamba. You do
not install PyTorch by hand on each host.
Pipeline
clone_repo (local, shell) git clone ──► repo/
│
├─► prepare_pretrain_val (CPU) Pile val shard ──► pile_dev.h5
├─► prepare_pretrain_train (CPU) Pile train shards ──► pile_train.h5
│ └─► train_pretrain (GPU) pretrain base model ──► base_pretrained.pt
│ └─► train_sft (GPU) SFT ──► sft.pt
│ └─► eval_sft GSM8K accuracy ──► stage_table.jsonl
│
├─► prepare_sft_data (CPU) Alpaca/Dolly/GSM8K ──► sft_packed.h5
├─► prepare_preference_data (CPU) HH-RLHF/UltraFeedback ─► preferences.jsonl
└─► prepare_rl_prompts (CPU) GSM8K + arithmetic ──► rl_prompts_*.jsonl
Inputs and outputs
Inputs
None. The workflow clones its own training repo. It downloads the Pile shard, Alpaca, Dolly, GSM8K, HH-RLHF, and UltraFeedback from Hugging Face.
Outputs land in workflow_results/:
ckpts/base_pretrained.pt: the pretrained base model.ckpts/sft.pt: the fine-tuned model.logs/pretrain.jsonl,logs/sft.jsonl: one JSON record per logged training step.logs/stage_table.jsonl: the GSM8K accuracy row for the SFT checkpoint.data/: the tokenized datasets, plus the preference and RL prompt sets.
Each checkpoint embeds its model config, training step, and metrics. You need no separate architecture file to load it.
Run the workflow
Install the horus-runtime and the plugins one time:
uv sync
If you do not have uv, install it first:
curl -LsSf https://astral.sh/uv/install.sh | sh
You can also install the packages with pip:
pip install horus-runtime horus-environments
Then run the workflow:
uv run horus run workflow.yaml
To set the model scale, edit --config in the train_pretrain and train_sft
commands. The default is configs/smoke/{pretrain,sft}.json: a tiny model
(n_embed=64, n_head=4, n_blocks=2) trained for about 10 to 20 steps. Point
it at configs/pretrain.json and configs/sft.json for the repo's default
~400M-parameter model.
To override a training field, add --field value to the task's command:. Any
field on PretrainConfig or SFTConfig works, for example --lr,
--batch_size, --train_steps, or --use_wandb true --wandb_project <name>.
To pull more pretraining text, raise --num_shards on prepare_pretrain_train.
The default is 1.
For a faster first run on a smaller dataset, start with the TinyStories workflow.
References
Run this workflow
The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.