October 5, 2026 · Temple Compute
Train an LLM From Scratch: A Pretrain, SFT, Eval Pipeline
Most "train an LLM from scratch" guides end with one long script. It downloads data, tokenizes it, trains, fine-tunes, and evaluates, top to bottom, on one machine. That works once. It stops working the first time SFT crashes an hour in and the only way back is to rerun tokenization, or to comment out half the file.
We added two workflows to the open-source library that do the same job as a pipeline instead:
- Train an LLM From Scratch: Pretrain, SFT and Eval pretrains on a shard of the Pile, fine-tunes on Alpaca, Dolly and GSM8K, and scores GSM8K accuracy.
- Tiny LLM From Scratch on TinyStories pretrains the same model on TinyStories and samples a story. It runs on a laptop in minutes.
Both use the
train-llm-from-scratch
repo: a decoder-only Transformer in pure PyTorch, with no transformers, trl,
or peft. The workflows add no new model code. They split the run into stages
with explicit inputs and outputs, and let Horus run them.
The stages
Each stage below is one task in the workflow YAML. Each writes a file the next stage reads.
Clone. A shell task shallow-clones the training repo. Every other task takes the checkout as an input, so the whole DAG has one root.
Data prep and tokenize. The data tasks stream a dataset and batch-tokenize it into a flat-token HDF5 file. In the full workflow there are five of them:
- Pile validation shard to
pile_dev.h5. - Pile training shards to
pile_train.h5.--num_shardssets how much text. - Alpaca, Dolly and GSM8K, rendered with a chat template and a loss mask, packed
to
sft_packed.h5andsft_dev_packed.h5. - Preference pairs from HH-RLHF and UltraFeedback, for a later reward model or DPO run.
- GSM8K and arithmetic prompt sets, for a later PPO or GRPO run.
The last two are not consumed yet. They sit in the DAG so a follow-up workflow
can pick them up. In the tiny workflow, a script bundled with the workflow
tokenizes TinyStories with tiktoken's r50k_base into the same HDF5 layout.
Pretrain. pretrain_base.py trains the base model on the tokenized train
and dev files and writes base_pretrained.pt, plus one JSON log record per
logged step. The checkpoint embeds its model config, training step, and
metrics, so later stages need no separate architecture file.
SFT. train_sft.py loads the base checkpoint and fine-tunes it on the
packed instruction rows. Output: sft.pt.
Eval. eval_post_training.py greedily decodes the SFT checkpoint on
held-out GSM8K questions and appends the accuracy to stage_table.jsonl.
Sample. The tiny workflow has no SFT or eval. Instead, chat.py --raw --greedy continues the prompt "Once upon a time," for 200 tokens and writes it
to a text file. It is a base-model continuation, a quick check that the model
learned something.
Out of the box, both workflows use the repo's smoke-scale configs: a tiny model
trained for a few dozen steps at most. Change --config to the repo's default
configs for its ~400M-parameter model.
Why this is a heterogeneous-compute DAG
Look at what each stage actually needs.
The clone and data stages are CPU and network work. They download, tokenize, and write HDF5. A GPU sits idle through all of it. They also don't depend on each other, so five of them can run at once.
Pretraining and SFT are the GPU stages. At smoke scale they finish in minutes on a laptop. At the default model size they need a real GPU and hours to days.
Eval and sampling load one checkpoint and decode. They need far less than training does.
So one LLM pretraining pipeline wants at least two kinds of machine. Run it all on a GPU node and you pay for the GPU while tokenizers run. Run it all on a CPU box and training never finishes. Split it by hand and you are copying HDF5 files and checkpoints between hosts, hoping each stage ran on the right input.
How Horus routes each stage
In Horus, every task declares its inputs, outputs, executor, and target:. The
edges: connect outputs to inputs. A task starts only once its upstream
artifacts exist.
The target: is per task. The shipped YAML runs every task locally, so it works
on a laptop with no setup beyond uv sync. To move training to a GPU machine,
change the target: of train_pretrain and train_sft. The data stages stay
on CPU. The runtime.command string stays the same. Horus moves the tokenized
data to the training host and brings the checkpoint back.
The executor builds the Python environment from the workflow's conda_env.yaml
with micromamba, so you do not install PyTorch by hand on each host.
Completed stages are skipped
Every stage is resumable. On a rerun, a task whose outputs already exist does not run again.
That changes how you iterate. Tune the SFT learning rate and rerun: only SFT and eval execute. The Pile shards, the packed SFT data, and the base checkpoint are reused. Change the sample prompt in the tiny workflow and only the sample stage runs. No commented-out lines, no re-tokenizing.
Start small with TinyStories, then scale
The Pile is roughly 900 GB. TinyStories is a couple hundred MB of short, simple stories. That makes the tiny workflow the right first run: it pretrains the same model on 20,000 stories and samples a continuation, on a laptop, in minutes.
Once it works, scale one knob at a time:
- Raise
--max_docson the TinyStories tokenize stages, or drop it to use all ~2.1M train stories. - Raise
--train_stepson the pretrain stage, for example--train_steps 500 --lr 3e-4. - Move to the full workflow for Pile pretraining, SFT, and GSM8K eval.
- Switch
--configto the repo's default model size and point the training stages at a GPU target.
The DAG is the same at every step. Only arguments and targets change.
FAQ
How do I train an LLM from scratch?
Tokenize a text corpus, pretrain a decoder-only Transformer on it, fine-tune the base model on instruction data, then evaluate it. The full workflow runs exactly those stages, in pure PyTorch, as separate tasks you can rerun one at a time.
Can I train an LLM from scratch on a laptop?
Yes, at small scale. The TinyStories workflow
uses a tiny config (n_embed=128, n_blocks=2, 20 training steps) and runs on
a laptop CPU or GPU in minutes. The repo's ~400M-parameter default needs a real
GPU and hours to days.
How do I run LLM pretraining on a GPU cluster without moving data by hand?
Keep data prep on CPU and set the target: of the pretrain and SFT tasks to
your GPU machine. Horus moves the tokenized data and the checkpoints between
stages.
What does a pretrain, SFT, eval pipeline output?
A base checkpoint, an SFT checkpoint, per-step JSONL training logs, and a GSM8K accuracy table. Each checkpoint carries its own model config.
Run it
Both workflows are in the open-source library. Clone, uv sync,
uv run horus run workflow.yaml. For how Horus compares to other tools for
mixed CPU and GPU pipelines, see
the best workflow orchestration tools for heterogeneous compute.
To run the training stages on your own GPUs or cloud from one place, open Horus.