All workflows

AI and ML

Tiny LLM From Scratch on TinyStories

Pretrain a tiny LLM from scratch on TinyStories in minutes on a laptop, then sample a story. Same code as the full pipeline, swap the target to scale up.

PyTorchtiktokendatasetsh5pyconda

What this workflow does

This workflow pretrains a small decoder-only Transformer from scratch on TinyStories. TinyStories is a set of short, simple, GPT-generated children's stories. It is a couple hundred MB, against roughly 900 GB for the Pile.

It uses the same model code as the full pretrain, SFT and eval workflow, from the train-llm-from-scratch repo. It skips instruction tuning and eval. The deliverable is a pretrained checkpoint and a sample generation that continues the prompt "Once upon a time,".

The goal is to see a from-scratch LLM learn something legible in minutes, on a laptop, with no large disk or bandwidth budget.

The compute problem

The pipeline is small, but its stages still differ.

The clone stage needs git and a few seconds.

The two tokenize stages download TinyStories and write flat-token HDF5 files. They are CPU and network work, and they run in parallel.

The pretrain stage is the only one that benefits from a GPU. At the shipped config it runs on a laptop CPU or GPU in minutes.

The sample stage loads the checkpoint and decodes 200 tokens.

The real problem shows up later. When you scale the model or the data, you want the same pipeline on a GPU machine. You do not want a second script for it.

How Horus solves it

Horus runs each stage as its own task with declared inputs and outputs. The edges: wire the HDF5 files into train_pretrain_tiny and the checkpoint into sample_story.

Each task has its own target:. The shipped YAML runs every task locally. To train on a GPU machine, change the target: of train_pretrain_tiny. The tokenize stages stay on CPU. The command does not change, and Horus moves the data and the checkpoint for you.

On a rerun, a stage whose outputs already exist is skipped. Change the sample prompt and only sample_story runs again.

Horus builds the Python environment from conda_env.yaml with micromamba.

Pipeline

clone_repo (local, shell)                git clone          ──► repo/
   │
   ├─► prepare_tinystories_train (CPU)   20,000 stories     ──► tinystories_train.h5
   ├─► prepare_tinystories_val   (CPU)   2,000 stories      ──► tinystories_val.h5
   │        └─► train_pretrain_tiny      pretrain           ──► tinystories_pretrained.pt
   │                 └─► sample_story    greedy continuation ─► tinystories_sample.txt

Inputs and outputs

Inputs

None. The workflow clones its own training repo and downloads TinyStories from Hugging Face. The tokenizer script, scripts/prepare_tinystories.py, ships with the workflow. The cloned repo's own data script is hardcoded to the Pile.

Outputs land in workflow_results/:

  • ckpts/tinystories_pretrained.pt: the pretrained checkpoint. It embeds its model config, training step, and metrics.
  • logs/pretrain.jsonl: one JSON record per logged training step.
  • samples/tinystories_sample.txt: a greedy, raw continuation of "Once upon a time,".

The sample is a base-model continuation. The workflow never runs SFT, so the model has no instruction-following behavior.

Run the workflow

Install the horus-runtime and the plugins one time:

uv sync

If you do not have uv, install it first:

curl -LsSf https://astral.sh/uv/install.sh | sh

You can also install the packages with pip:

pip install horus-runtime horus-environments

Then run the workflow:

uv run horus run workflow.yaml

To set the dataset size, edit --max_docs on the two tokenize stages. The defaults are 20,000 train and 2,000 validation stories. Full TinyStories has about 2.1M train stories. Drop the flag to use all of them.

To set the model scale, edit the train_pretrain_tiny command. It loads configs/smoke/pretrain.json: n_embed=128, n_head=4, n_blocks=2, context_length=256, train_steps=20. Override any PretrainConfig field with --field value, for example --train_steps 500 --lr 3e-4.

To change the sample, edit --prompt, --max_new_tokens, --temperature, or --top_p on sample_story.

Tokens use tiktoken's r50k_base encoding (50,257 tokens). This matches the repo's padded vocab_size=50304, so there is no vocab mismatch.

References

Run this workflow

The workflow is open source. Clone the pantheon repository and run it with the horus-runtime engine. To run it on managed compute without a cluster of your own, open it in Temple Compute OS.