All posts

October 5, 2026 · Temple Compute

Train an LLM From Scratch: A Pretrain, SFT, Eval Pipeline

Most "train an LLM from scratch" guides end with one long script. It downloads data, tokenizes it, trains, fine-tunes, and evaluates, top to bottom, on one machine. That works once. It stops working the first time SFT crashes an hour in and the only way back is to rerun tokenization, or to comment out half the file.

We added two workflows to the open-source library that do the same job as a pipeline instead:

Both use the train-llm-from-scratch repo: a decoder-only Transformer in pure PyTorch, with no transformers, trl, or peft. The workflows add no new model code. They split the run into stages with explicit inputs and outputs, and let Horus run them.

The stages

Each stage below is one task in the workflow YAML. Each writes a file the next stage reads.

Clone. A shell task shallow-clones the training repo. Every other task takes the checkout as an input, so the whole DAG has one root.

Data prep and tokenize. The data tasks stream a dataset and batch-tokenize it into a flat-token HDF5 file. In the full workflow there are five of them:

  • Pile validation shard to pile_dev.h5.
  • Pile training shards to pile_train.h5. --num_shards sets how much text.
  • Alpaca, Dolly and GSM8K, rendered with a chat template and a loss mask, packed to sft_packed.h5 and sft_dev_packed.h5.
  • Preference pairs from HH-RLHF and UltraFeedback, for a later reward model or DPO run.
  • GSM8K and arithmetic prompt sets, for a later PPO or GRPO run.

The last two are not consumed yet. They sit in the DAG so a follow-up workflow can pick them up. In the tiny workflow, a script bundled with the workflow tokenizes TinyStories with tiktoken's r50k_base into the same HDF5 layout.

Pretrain. pretrain_base.py trains the base model on the tokenized train and dev files and writes base_pretrained.pt, plus one JSON log record per logged step. The checkpoint embeds its model config, training step, and metrics, so later stages need no separate architecture file.

SFT. train_sft.py loads the base checkpoint and fine-tunes it on the packed instruction rows. Output: sft.pt.

Eval. eval_post_training.py greedily decodes the SFT checkpoint on held-out GSM8K questions and appends the accuracy to stage_table.jsonl.

Sample. The tiny workflow has no SFT or eval. Instead, chat.py --raw --greedy continues the prompt "Once upon a time," for 200 tokens and writes it to a text file. It is a base-model continuation, a quick check that the model learned something.

Out of the box, both workflows use the repo's smoke-scale configs: a tiny model trained for a few dozen steps at most. Change --config to the repo's default configs for its ~400M-parameter model.

Why this is a heterogeneous-compute DAG

Look at what each stage actually needs.

The clone and data stages are CPU and network work. They download, tokenize, and write HDF5. A GPU sits idle through all of it. They also don't depend on each other, so five of them can run at once.

Pretraining and SFT are the GPU stages. At smoke scale they finish in minutes on a laptop. At the default model size they need a real GPU and hours to days.

Eval and sampling load one checkpoint and decode. They need far less than training does.

So one LLM pretraining pipeline wants at least two kinds of machine. Run it all on a GPU node and you pay for the GPU while tokenizers run. Run it all on a CPU box and training never finishes. Split it by hand and you are copying HDF5 files and checkpoints between hosts, hoping each stage ran on the right input.

How Horus routes each stage

In Horus, every task declares its inputs, outputs, executor, and target:. The edges: connect outputs to inputs. A task starts only once its upstream artifacts exist.

The target: is per task. The shipped YAML runs every task locally, so it works on a laptop with no setup beyond uv sync. To move training to a GPU machine, change the target: of train_pretrain and train_sft. The data stages stay on CPU. The runtime.command string stays the same. Horus moves the tokenized data to the training host and brings the checkpoint back.

The executor builds the Python environment from the workflow's conda_env.yaml with micromamba, so you do not install PyTorch by hand on each host.

Completed stages are skipped

Every stage is resumable. On a rerun, a task whose outputs already exist does not run again.

That changes how you iterate. Tune the SFT learning rate and rerun: only SFT and eval execute. The Pile shards, the packed SFT data, and the base checkpoint are reused. Change the sample prompt in the tiny workflow and only the sample stage runs. No commented-out lines, no re-tokenizing.

Start small with TinyStories, then scale

The Pile is roughly 900 GB. TinyStories is a couple hundred MB of short, simple stories. That makes the tiny workflow the right first run: it pretrains the same model on 20,000 stories and samples a continuation, on a laptop, in minutes.

Once it works, scale one knob at a time:

  1. Raise --max_docs on the TinyStories tokenize stages, or drop it to use all ~2.1M train stories.
  2. Raise --train_steps on the pretrain stage, for example --train_steps 500 --lr 3e-4.
  3. Move to the full workflow for Pile pretraining, SFT, and GSM8K eval.
  4. Switch --config to the repo's default model size and point the training stages at a GPU target.

The DAG is the same at every step. Only arguments and targets change.

FAQ

How do I train an LLM from scratch?

Tokenize a text corpus, pretrain a decoder-only Transformer on it, fine-tune the base model on instruction data, then evaluate it. The full workflow runs exactly those stages, in pure PyTorch, as separate tasks you can rerun one at a time.

Can I train an LLM from scratch on a laptop?

Yes, at small scale. The TinyStories workflow uses a tiny config (n_embed=128, n_blocks=2, 20 training steps) and runs on a laptop CPU or GPU in minutes. The repo's ~400M-parameter default needs a real GPU and hours to days.

How do I run LLM pretraining on a GPU cluster without moving data by hand?

Keep data prep on CPU and set the target: of the pretrain and SFT tasks to your GPU machine. Horus moves the tokenized data and the checkpoints between stages.

What does a pretrain, SFT, eval pipeline output?

A base checkpoint, an SFT checkpoint, per-step JSONL training logs, and a GSM8K accuracy table. Each checkpoint carries its own model config.

Run it

Both workflows are in the open-source library. Clone, uv sync, uv run horus run workflow.yaml. For how Horus compares to other tools for mixed CPU and GPU pipelines, see the best workflow orchestration tools for heterogeneous compute.

To run the training stages on your own GPUs or cloud from one place, open Horus.