Training a 58M language model from scratch on free GPUs

Two days, zero dollars, and a full public record: micro-ablations before full runs, an optimizer that won 7 of 7 checkpoints and was still not promoted, a validation plateau resolved by a deterministic evaluation, and the quota exhaustion that ended training at 73% of the schedule.

The goal and the constraint

Dedicated page: KALIA 0.1.2The full paper-style record — model specification, evaluation, a replayable training console, and every decision and incident.

I trained KALIA, a 58M-parameter language model, from random initialization to coherent story generation in two days, using only free-tier Kaggle GPUs. No pretrained weights, no distillation, no fine-tuning, and zero dollars of compute. It writes short stories, scores 61.4% on PIQA, and every weight in it exists nowhere else.

Two rules shaped the project. The first was from-scratch: fine-tuning forks someone else's brain, and I wanted an artifact where every byte was accountable — the corpus mixture, the tokenizer, the architecture, the optimizer, the exact run. The second rule was verifiability: every experiment is pre-registered with a SHA-256 hash before it runs, every decision is numbered, and every incident is published, including the ones that make me look bad.

  • The constraint stack: 30 GPU-hours per week on 2x NVIDIA T4 (32GB total).
  • 8.5-hour maximum session length, so the trainer had to be resumable by design.
  • API-triggered runs cannot read secrets, so real training sessions start from the browser.
  • Total compute spent on the released model: about 20 GPU-hours.

Measure before you spend

When quota is the scarce resource, the worst thing you can do is discover a bad recipe with it. Every change is first screened at 30M parameters on 50M tokens, roughly 30 minutes per arm, with identical seed and token budget. Only winners are promoted to full runs.

# Screen three optimizers at 30M params, 500 steps each (~90 minutes total)
python ablate.py --arms micro-base,micro-muon,micro-muon-qk --steps 500
The micro-ablation ladder: one change per arm, same data and seed.
  • AdamW: 3.8041 step-700 validation loss.
  • Muon: 3.5937, a 0.21 improvement.
  • Muon + QK-Norm + logit soft-capping: 3.5103, a 0.29 improvement.
  • The ordering held at every evaluation checkpoint, so it was not noise.
Bar chart comparing step-700 validation loss for AdamW, Muon, and Muon with QK-Norm and logit soft-capping; lower is better, and the combined recipe wins.
The optimizer ladder: each change screened at 30M params before spending full-run quota.

The full-scale run confirmed it: the new recipe reached the AdamW baseline's final loss with about 23% fewer tokens, and kept a persistent gap of roughly 0.15 nats at equal step counts.

Muon+ validated, not promoted

Muon+ adds one post-polar normalization step to Muon, and the papers report gains from 60M parameters upward. In our ablation it beat plain Muon on 7 of 7 checkpoints. It was not promoted to the full-scale recipe.

The gain was 0.015 nats. The promotion threshold, fixed before the experiment ran, was 0.02. A threshold that is bent for a favored result is not a threshold, so the full-scale model kept plain Muon and the 0.015 result remains in the record as validated but below the bar.

Architecture search: three negatives and one finding

With the optimizer settled, the next question was shape. Four arms, same data, same seed, 763 steps: the control design, a looped model that passes through the same weights twice for double effective depth, a thin-and-deep model, and grouped-query attention.

  • Control (Muon + QK-Norm): 3.4924 step-700 validation loss.
  • Looped depth: 3.5012 — quality-neutral.
  • Thin and deep: 3.6969 — clearly worse.
  • Grouped-query attention: 3.5031 — quality-neutral, with 4% fewer parameters.
Bar chart comparing step-700 validation loss for the control architecture, looped depth, thin-and-deep, and grouped-query attention; the control arm is lowest.
Nothing beat the control arm by the pre-set margin, so the negatives shipped with the results.

Nothing beat control by the pre-set margin, so the architecture stayed as it was and the negatives were published. One finding survived anyway: the looped model ran only 1.35x slower per step, not 2x, because the reused weights stay hot in cache between passes. Double effective depth for a third more compute is interesting economics — it just did not buy quality at this scale.

The run: a plateau, a decay, and a quota wall

The released model trained across three sessions. For a thousand steps the validation loss refused to move while training loss kept falling — the classic shape of a noisy eval or a real plateau, and impossible to tell apart from a single point.

  • Step 1500: 2.5270, then 1750: 2.5453, 2000: 2.6334, 2250: 2.6261, 2500: 2.5736 — flat inside a band of about 0.1.
  • Step 2750: 2.4438, then 3000: 2.3986 as the cosine decay bit, then 3250: 2.5138 — the swings were the eval, not the model.
  • The weekly GPU quota ran out at step 3478 of 4770, 73% of the schedule.
Line chart of validation loss from step 250 to 3250: a steep descent from 4.0881 to 2.5270 by step 1500, then a plateau band between 2.40 and 2.63 for the remaining 1,750 steps, plus a separate deterministic evaluation point at step 3478 with loss 2.4366.
The full curve: the descent finishes by step 1,500 and everything after it oscillates inside 0.23 nats. The first half was missing from the published log — a mid-run code change dropped it — and was rebuilt from the model repository commit history.

A deterministic evaluation settled the question the noisy training evals could not: 100 fixed-seed batches over 819,200 tokens returned 2.4366 loss and 0.8184 bits-per-byte, on the original corpus's held-out set. That set was later rebuilt for licence compliance and the rebuilt one is 0.62 nats harder for identical weights, so the figure is not comparable to anything measured afterwards — a lesson that arrived as incident D43.

Evaluation beyond loss

  • Probe held-out loss on 20 fixed sentences: 3.2303, or 0.9415 bits-per-byte.
  • Zero-shot benchmarks, 500 samples each: PIQA 61.4%, ARC-Easy 45.8%, HellaSwag 36.8% (normalized), WinoGrande 50.2%, LAMBADA 23.0% accuracy at perplexity 194.
  • For scale context, leaderboard tables list OPT-125M — twice the parameters and roughly 160x the training tokens — at PIQA 63.0% and ARC-Easy 43.5%. Treat that as context, not a head-to-head: harness versions differ.
Bar chart of zero-shot benchmark accuracy: PIQA 61.4, ARC-Easy 45.8, HellaSwag 36.8, WinoGrande 50.2, with dashed chance lines at 50 and 25 percent.
Above chance on every task; genuinely competitive on PIQA and ARC-Easy for a 58M storyteller.

The metric I am most attached to is custom. Take held-out text, measure the model's loss on it forward, then measure its loss on the same text with the tokens reversed. Forward 3.23, reversed 9.29. The Abhimanyu gap is the difference: 6.06 nats. The model can enter fluent text but cannot exit it — the computational form of the warrior who entered the Chakravyuha formation and could not find his way out. Random guessing would be about 10.8, so reversed text is nearly as foreign to the model as noise.

# Reproduce the evaluation on the released checkpoint
python eval_probes.py --ckpt ckpt.pt --out out/eval/report.md
python eval_reversibility.py --ckpt ckpt.pt --out out/eval/reversibility.md
python eval_val.py --ckpt ckpt.pt --val-bin val.bin --batches 100
The evaluation suite runs on CPU; no GPU quota is spent.

The record is the point

Any training run produces a loss curve. What makes this one auditable is everything around it: 43 numbered decisions, 14 published incidents, three hash-anchored pre-registrations (plus two amendments), and a dated journal. The incidents are the useful part.

  • A resume race where both distributed workers wrote the same checkpoint file; fixed with a single downloader, an atomic swap, and a barrier.
  • A silent success: a failed training subprocess was still marked complete, because shell-style commands do not fail notebook cells; fixed by asserting exit codes.
  • A misreported duration: I described a 44-minute ablation as having run five hours, because I trusted my sense of time instead of the run-start timestamp. The correction is in the journal.
  • A documentation error caught late: our own docs described the released model as Muon+ when the config proved it was plain Muon. Every public text was corrected, and the hashed v0.2.0 pre-registration received a registered amendment rather than a silent edit.

None of this is glamorous. All of it is why the numbers in this post can be checked by anyone with a browser, and why I trust them myself.

What is public now

  • Weights, model card, configs, and the full resumable checkpoint on HuggingFace (Apache-2.0).
  • Source, tests, notebooks, journal, decisions, incidents, and the pre-registration ledger on GitHub (MIT).
  • Training data is never redistributed; every source is attributed in the model card.
  • The Kaggle notebooks that built and evaluated the model are private for now, and the self-contained ones will be published next.
# Generate from the released checkpoint (CPU is fine; the model is ~230MB)
from huggingface_hub import hf_hub_download
path = hf_hub_download("kalia-lm/kalia-v012", "checkpoints/ckpt.pt")
# then, from a clone of the repository:
# python sample.py --ckpt <path> --prompt "Once upon a time"
The raw checkpoint loads with the project’s own model code, no transformers required.

Reproduce it

git clone https://github.com/subhajitlucky/kalia && cd kalia
python -m venv .venv && . .venv/bin/activate && pip install -r requirements.txt
python -m pytest tests/ -v          # 64 tests, all green
The full test suite runs on CPU in about 16 seconds.

Data preparation, training, and evaluation all run on free Kaggle notebooks, and every one of them is in the repository. The pre-registration hashes are in the ledger, so anyone can verify that a prediction existed before its result did.

What is next

The Abhimanyu gap was the thread I wanted to pull. The pre-registered experiment tested whether chunk-preserving reversal training — reversing the order of short chunks while keeping tokens inside each chunk readable — closes the gap without hurting forward loss. The prediction, threshold, and analysis plan were hashed before the run. It failed: the gap came out 0.14 nats worse than the control, on both seeds, on both metrics, and the transform was rejected by the rule that had already decided what would count as success.

Then came v0.2.0, now training: a compliance-clean corpus of 2.4B tokens, with only those changes that pass promotion rules hashed before any result existed. The reversal transform was not one of them. Its final evaluation — yardstick, thresholds, and the claims it is forbidden to make — was registered while it was still training.