open model release · 2026-09-23 · $0 compute

KALIA

अथ शब्दानुशासनम् — “Now begins the discipline of words.”

This page documents a 58M-parameter language model trained from random initialization to coherent story generation in two days, on free-tier Kaggle GPUs, at zero compute cost. Section 1 fixes the model specification; Section 2 reports the evaluation; Sections 3 and 4 present unedited samples and an interactive replay of the training log. The appendices record the day-by-day log, all 43 numbered decisions, and all 15 incidents. Every experiment was pre-registered before it ran, and no claim on this page is a screenshot — each links to the primary log.

parameters57.9M
tokens seen1.82B
val loss2.4366
bits/byte0.8184
PIQA61.4%
Abhimanyu gap6.06 nats
GPU-hours~20
cost$0

1. Model specification

Table 1 · frozen by pre-registration

A decoder-only transformer trained from random initialization — no pretrained weights, no distillation, no fine-tuning. Everything below was fixed before the final run started, and the next version may only change what passes a hashed promotion rule.

Parameters57,856,256
Shape10 layers x 512 dim x 8 heads
Context1,024 tokens
Vocabulary50,257 (GPT-2 BPE)
OptimizerMuon 0.02 + AdamW 6e-4
Schedulecosine, 500-step warmup, 4,770 planned steps
Batch524,288 tokens per step (8 x 32 x 1024 x 2 GPUs)
Precisionfp16 with gradient scaling, DDP across 2x T4
CorpusTinyStories + FineWeb-Edu, 2.46B tokens
Tokens seen1.82B (stopped at step 3,478 of 4,770)
Computeabout 20 free-tier GPU-hours
Cost$0

2. Two checkpoints, one yardstick

Table 2 · deterministic 100-batch evaluation, 819,200 tokens, fixed seed

The corpus was rebuilt for licence compliance, which replaced a held-out set and made it 0.62 nats harder for identical weights. So v0.1.2 was re-measured on the rebuilt set before anything could be compared, and the older figure of 2.4366 is never tabulated next to a rebuilt one. Both versions below are scored the same way, by the same code path, on the same data.

KALIA 0.1.2 releasedval 3.0533 · 0.9992 bpB · 1.82B tokens · stopped at 73% of its schedule; the first public release
KALIA 0.2.0 experimentval 2.8248 · 0.9244 bpB · 2.50B tokens · same recipe, compliance-rebuilt corpus; full 4,770-step schedule

That is 0.2285 nats better — 4.6× the pre-registered 0.05 bar — from changing nothing but the data.

Grouped bars for five zero-shot tasks, both versions side by side, each over a grey band showing plus or minus one standard error of about two points. PIQA, HellaSwag and WinoGrande move up but sit inside their own error bands; LAMBADA and ARC-Easy move down by more than one band.
Both versions, same harness, same 500 samples, same code path. The grey band is the measurement’s own uncertainty, so a reader can see which bars actually separate and which merely look like they do.

The same five tasks as exact numbers

PIQA63.8%was 61.4 · +2.40 (1.1σ) · physical commonsense
HellaSwag39.8%was 36.8 · +3.00 (1.5σ) · sentence completion (normalized)
WinoGrande50.8%was 50.2 · +0.60 (inside noise) · coreference
LAMBADA20.8%was 23.0 · -2.20 (1.2σ) · long-range cloze
ARC-Easy42.0%was 45.8 · -3.80 (1.7σ) · science questions

Read that table with the error bars attached. At 500 samples these benchmarks carry a standard error of roughly 2 points, so three of the five movements are smaller than the measurement’s own noise and cannot be called improvements. The two that fell — ARC-Easy at 1.7σ and LAMBADA at 1.2σ — are the honest signal, and they are why v0.2.0 is published as an experiment rather than as a release.

The pre-registered rule was that no task may regress by more than 1.0 point. v0.2.0 fails it on two. The thresholds were not moved afterwards, and the noise floor was published beside the verdict rather than used to rescue it.

3. Where the tokens came from

Table 3 · 2.4B tokens

FineWeb-Edu (dedup) · 60%ODC-By-1.0 · educational web text
TinyStories · 20%CDLA-Sharing-1.0 · children's stories; almost none long enough to matter
Cosmopedia v2 · 15%Apache-2.0 · synthetic textbooks
Python (per-file filtered) · 5%7 permissive licences · the earlier unfiltered slice was 41.7% copyleft
The training mixture as a stacked bar, beside the share of each source whose documents are at least as long as the 1,024-token context: 26.2% for FineWeb-Edu, 0.03% for TinyStories, 13.2% for Cosmopedia, and 17.7% for the mixture overall.
The mixture on the left looks balanced. The panel on the right is why it isn’t: a fifth of the tokens supply 0.03% of the documents long enough for the model to see even one whole document, and only 17.7% of the corpus clears the bar at all. That number is invisible in a table of proportions.

Only 17.7% of training tokens sit inside a document at least as long as the 1024-token context. TinyStories — a fifth of the mixture — supplies 0.03% of them. LAMBADA asks a model to hold a discourse and recall its final word, so its regression is the predicted direction for this corpus rather than a mystery. The same probe found a candidate source at 99.3%, which makes the remedy measured rather than guessed.

One thing was found here that nobody was looking for. The Python slice of the previous corpus was built without a licence check, and sampling 20,000 source files measured 41.7% of characters under copyleft licences — about 2% of that training set. The filter that prevents this existed, was tested, and was simply written 40 minutes after the corpus that needed it. It is disclosed on the released model card rather than quietly fixed.

4. Evaluation — KALIA 0.1.2, first release

Table 2 · 0-shot, 500 samples, lm-evaluation-harness

PIQA61.4%physical commonsense
ARC-Easy45.8%science questions
HellaSwag36.8%sentence completion (normalized)
WinoGrande50.2%coreference
LAMBADA23.0%long-range cloze; perplexity 194

Held-out loss on the deterministic 100-batch evaluation is 2.4366 (819,200 tokens, fixed seed) — 0.8184 bits-per-byte, measured on the original corpus’s held-out set. That set was later rebuilt for licence compliance; the rebuilt one is 0.62 nats harder for these same weights (3.0533), which is why v0.2.0 is scored against the rebuilt set and the two figures are never tabulated together. A 58M storyteller at 61.4% PIQA and 45.8% ARC-Easy is competitive with 125M-class models trained on roughly 160x more tokens, and weakest on the one task that requires long-range narrative memory (LAMBADA, 23%). The custom metric is the one that matters next: the Abhimanyu gap — reversed-text loss (9.29) minus forward loss (3.23) — is 6.06 nats. The model can enter fluent text but cannot exit it.

The pre-registered experiment that tested whether training on chunk-preserving reversal closes that gap has now run, and it failed — reported here with the same prominence a win would have had. At 30M parameters, control models showed a 5.11-nat gap (confirming the effect is architectural, not an artifact of scale), and the treatment trained on 50% reversed chunks made it worse: gap 5.25 (+0.14) and forward loss +0.08, missing both pre-registered bars on both seeds. Reversing chunk order while preserving intra-chunk order does not teach token-level reversal. The transform was rejected by the hashed promotion rule and never shipped, and a stronger form is now a separate hypothesis rather than a quiet retry.

5. Unedited samples

temperature 0.8, top-k 200, checkpoint step 3,478

prompt: The little fox

The little fox was so restless. He wanted to explore the pond, but he was too small. Suddenly, he heard a voice. It said, "Do not go near the water". The fox looked up and saw a little girl smiling at him. The girl said, "Let's go have some fun the next time we get closer!".

Coherent scene, dialogue, and a promise of more — at 58M parameters.

prompt: The brave rabbit

The brave rabbit saw two and three friends flying together in the sky. They were laughing and exploring together. The friends sang, laughed and sang. They were so happy to have a friend like the brave seeker.

Repetition at "two and three friends" — the model's small vocabulary of ideas shows here.

prompt: A curious cat named

A curious cat named Max. One day, Tim went to the pond to catch some fish to eat. The fish was not happy and swam away. Tim saw a man fishing in a nearby river.

The honest failure mode: Max becomes Tim within two sentences. Entity drift is the reason the next experiment exists.

6. The run, replayed

Figure 1 · replayed from the primary log

The console replays the surviving training log — session three, steps 1,740 to 3,470, at every tenth step. Earlier sessions' rows were lost to a resume incident (I8, Appendix C). Blue markers are validation evaluations; the square is the deterministic evaluation that established the model had not regressed.

kalia-train-v012 · session 3 replay, steps 1,740–3,4701,740 / 4,770
loss3.1018
lr4.95e-4
tokens912,261,120
tok/s0

Appendix A. Day-by-day log

times in UTC, from the kernel run logs

The plan changed several times as measurements arrived. The entries below are the real order of events — each algorithm measured before the next was attempted — with the three decisive comparisons shown as figures.

Day 1 2026-09-22

  1. A repository and a rule

    The first commit lands at 06:30 UTC: a design document, an implementation plan, and a scaffold. The rule is fixed from the start — no pretrained weights, ever. The architecture is a small decoder-only transformer (10 layers, 512 dimensions, RoPE, RMSNorm, SwiGLU) chosen to fit two T4s without gradient checkpointing. The suite grows test-first from the first minute.

  2. Tokenizing on free CPU

    Data preparation runs on a Kaggle CPU notebook, which costs no GPU quota. TinyStories and FineWeb-Edu become 2.46B tokens of uint16 shards in about 37 minutes: 473,992,236 tokens of stories and 2,000,001,223 tokens of educational web text.

  3. First GPU run: out-of-memory at step one

    The first training run fails at step one with an OutOfMemoryError. Batch 32 at context 1024 does not fit a T4. The fix is the standard one — micro-batch 8 with 32 gradient-accumulation steps, the same effective batch — plus expandable memory segments. Three further incidents land in the same hour: API runs cannot read secrets, the dataset uploader drops subdirectories, and a failed subprocess is silently marked complete. All four are fixed and recorded.

  4. Micro-ablations before full runs

    Micro-ablations run at 30M parameters on 50M tokens, about 30 minutes per arm. AdamW scores 3.8041 step-700 validation loss; Muon scores 3.5937; Muon with QK-Norm and logit soft-capping scores 3.5103. The ordering holds at every checkpoint, so it is not noise. Only the winner receives full-scale quota.

    Bar chart of step-700 validation loss for AdamW (3.8041), Muon (3.5937), and Muon with QK-Norm and logit soft-capping (3.5103); lower is better.
    Figure A1 — the optimizer screen: each change measured at 30M parameters before any full run.
  5. Two full runs launch

    The AdamW baseline and the new Muon recipe train in parallel sessions. The baseline finishes its first session at step 2,250 with validation loss 3.2702. The Muon recipe ends at step 1,738 at 3.2214 — already past the baseline final loss with roughly 23% fewer tokens.

Day 2 2026-09-23

  1. Resume, after fixing the race

    Overnight, session two crashed on resume: both distributed workers wrote the same checkpoint file and one read it mid-write. The fix is a single downloader, an atomic file swap, and a barrier. Session three resumes cleanly and runs for eight and a half hours.

  2. Muon+ validated, not promoted

    Muon+ — one post-polar normalization step — beats plain Muon on 7 of 7 checkpoints, by 0.015 nats. The promotion threshold, fixed before the run, was 0.02. It is not promoted. The full-scale model keeps plain Muon, and the result is recorded as validated but below the bar.

  3. The learning-rate sweep closes

    Four arms: 0.015 scores 3.4943, 0.02 scores 3.4941, 0.03 scores 3.5027, 0.06 scores 3.5380. The recipe needed no change. Hyperparameter tuning is now closed — there is no headroom left to buy.

  4. The v2 corpus, then the compliance rebuild

    A new mixture is built for v0.2.0: FineWeb-Edu, TinyStories, Cosmopedia, and permissively licensed Python. Git history shows why the first code shard was never filtered — the corpus that built v0.1.2 was created 40 minutes before the per-file license filter existed, so the shard streamed with no licence check at all. It is rebuilt with that filter — MIT, Apache, BSD, ISC, Unlicense, CC0 only. Nothing unclear-licensed will ever train a public model.

  5. Architecture ablation

    Four arms, same data, same seed: control, looped depth (the same weights applied twice), thin-and-deep, and grouped-query attention. Control wins at 3.4924. Looped and GQA are quality-neutral; thin-and-deep loses 0.2 nats. One finding survives: the looped model runs 1.35x slower per step, not 2x, because reused weights stay hot in cache.

    Bar chart of step-700 validation loss for the control architecture (3.4924), looped depth (3.5012), thin-and-deep (3.6969), and grouped-query attention (3.5031).
    Figure A2 — the architecture screen: no variant beats control by the pre-set 0.02 margin.
  6. A validation plateau

    Validation loss has not moved for a thousand steps — flat between 2.53 and 2.63 — while training loss keeps falling. A noisy evaluation and a real plateau are indistinguishable from a single point, so both possibilities are recorded and the run continues.

    Line chart of validation loss from step 250 to 3250: a steep descent from 4.0881 to 2.5270 by step 1500, then a plateau band between 2.40 and 2.63 for the remaining 1,750 steps, and a separate deterministic evaluation point at step 3478 with loss 2.4366.
    Figure A3 — the full curve: the descent finishes by step 1,500, and everything after it oscillates inside 0.23 nats. The first half was recovered from the model repository commit history (incident I14).
  7. Quota exhausted

    Two events land in the same minute. The architecture verdict: nothing beats control by the pre-set margin, so the recipe is unchanged. Then the quota push fails — the week’s 30 GPU-hours are exhausted. Experiments stop and the plan is revised.

  8. Session three ends at 73%

    The training session reaches step 3,478 of 4,770 — 73% of the cosine schedule — and stops at the session limit. The last logged validation eval was noisy (2.5138), so the question stands: did the model regress, or is the eval noisy?

  9. The deterministic eval

    A CPU evaluation — free, no GPU quota — runs 100 fixed-seed batches over 819,200 tokens: 2.4366 loss, 0.8184 bits-per-byte, on the held-out set of the corpus as it then stood. The model never regressed; the swings were sampling noise. Under the pre-registered stopping rule, the plateau plus the quota wall make the stop final at step 3,478.

  10. Evaluation on CPU

    Zero-shot benchmarks via lm-evaluation-harness: PIQA 61.4%, ARC-Easy 45.8%, HellaSwag 36.8%, WinoGrande 50.2%, LAMBADA 23.0% accuracy at 194 perplexity. And the custom metric: the Abhimanyu gap — reversed-text loss minus forward loss — is 6.06 nats. The model can enter fluent text but cannot exit it.

  11. Release

    Decision D41: stop, document the plateau, save the remaining quota for v0.2.0. Within the hour the repository and the model weights are public: 41 numbered decisions, 12 published incidents, two hash-anchored pre-registrations, and every evaluation log. Total compute: about 20 GPU-hours on the free tier. Cost: zero dollars.

Day 3 2026-09-27

  1. The experiment that was supposed to work

    X16 tests the one hypothesis that could explain the 6.06-nat gap: that a model can only learn to exit a sequence, never to enter one reversed. Four arms, two seeds, 500 steps each — control against 50% chunk-preserving reversal training. The prediction was registered with a hash before the run: close the gap by at least 0.10 nats, at no more than 0.02 nats of forward-loss cost.

  2. It failed

    Control gap 5.11 nats, treatment 5.25 — the gap got 0.14 nats worse, and forward loss rose 0.08. Both bars missed, on both seeds, in the same direction. The prediction that did hold is the more useful one: the gap is already 5.1 nats at 30M parameters, so it is a property of the architecture, not a symptom of undertraining.

  3. The published loss curve was missing its first half

    The training log on Hugging Face began at step 1,740, because a fix that restores logs on resume landed one session after that session had already started with the old code. 347 logged steps — the whole descent from 10.7 down to the plateau — were gone from the public record. Rebuilt from the repository’s own commit history, checked for gaps, and republished. The recovered curve supports the stop decision: the descent finishes by step 1,500 and the remaining 1,750 oscillate inside noise.

  4. The rebuilt corpus quietly replaced the measuring stick

    The new run’s first validation point reads 0.74 nats worse than the previous version’s while its training loss is better, which looks exactly like a catastrophic regression. It is not. The corpus rebuild replaced the held-out set, and the old weights score 2.3915 on the old set and 3.1070 on the new one — the set alone explains the entire gap. Both versions were re-measured on the rebuilt set so that anything crossing the rebuild boundary is believed only after being checked on both sides. A yardstick that moves silently is indistinguishable from a model that does.

    Validation loss for v0.2.0 falling from 4.83 to 2.88 over 4,770 steps, flat from step 3,250, with deterministic reference points of 3.0533 for v0.1.2 and 2.8248 for v0.2.0 on the same rebuilt yardstick.
    The answer to that question, once both versions were re-measured on the same set. The shaded band is the 1,500 steps at the end that bought nothing measurable.
  5. The rule does the deciding

    The promotion rule was hashed before the results existed, so there is nothing to renegotiate: the transform is rejected and never enters v0.2.0 (decision D42). A 50% reversed training mixture is not the same problem as reversing tokens, and the honest reading is that the chunk-order transform teaches nothing about entry into a reversed sequence. Token-level reversal stays on the list as a separate hypothesis, not a quiet retry.

Day 4 2026-09-28

  1. v0.2.0 starts

    The next run begins on a compliance-clean corpus rebuilt shard by shard (2.4B tokens; the code slice is now filtered to permissive licences only) with the frozen recipe, since the rejected experiment changed nothing about the recipe. Checkpoints sync to Hugging Face every 30 minutes so sessions can be interrupted safely.

  2. The benchmarks disagree with the loss curve

    An evaluation at the step-3,470 checkpoint splits the two signals we had trusted to agree. Deterministic loss improves by 0.169 nats on identical weights and an identical protocol, but ARC-Easy falls 4.6 points and LAMBADA 4.6. Read alone this looks like a catastrophic corpus regression. It is a mid-run reading, recorded here precisely because it turned out to be the wrong one.

  3. A frontier architecture arm

    Two 2026 changes screened against control at 30M parameters: rotary embeddings dropped on some layers, and a gated residual stream borrowed from a 125B model. Both pre-registered with hashed thresholds before the run, including an honest prior that the more interesting one was the likelier to be noise. Three arms, 500 steps, one seed.

  4. The best result in the project was one seed

    That gated residual looked like a win, so it was pushed further — a static learned per-channel modulation that never reads its input. It scored 4.6304 against control 4.7662, beating the data-dependent gate it was supposed to be a cheaper cousin of, and clearing its pre-registered bar by 0.1240. It was the largest effect any architecture arm had produced here, and it was real for exactly as long as it took to check. Replicated at two fresh seeds it inverted to +0.0516 — worse than control — so the sign flipped, not just the size.

    Paired bar chart. The static gate scores 4.6304 at seed 1337 but 4.9623 and 4.9087 at seeds 1338 and 1339, while control stays near 4.77 to 4.89. The effect reverses sign.
    Figure A3 — X20 and its replication. The bar that looked like the project’s best result.
  5. We had never measured the baseline

    The replication also returned a number the project did not have: how much the control moves when nothing about it changes. Two independent sessions at seed 1337 agree to 0.0018, so the machine is not the variable — but fresh seeds come in about 0.11 nats higher. That spread is larger than the entire −0.0436 the gated residual was credited with, and larger than the branch-norm null by an order of magnitude. Every single-seed delta in this project was read against a baseline nobody had measured. They were all noise, and the ordering between them was an artefact of ranking noise against noise.

  6. 41.7% of the code corpus is copyleft

    Twenty thousand files, 197 million characters, eleven minutes of CPU. 58.1% permissive, 41.7% not, and GPL-family licences alone cover 39.5% of all files. The published card had been listing three of the four training sources and omitting the code slice entirely — the one dataset that needed disclosing. The interpretation had been fixed before the number was known: a material share means the rebuild is the remediation, and the rebuild was already under way.

  7. Windows were crossing documents

    Auditing the data path finds the corpus is EOS-delimited and training never used that fact. Documents run about 200 tokens against a 1,024-token context, so nearly every window crossed a boundary with unmasked attention, and 0.3% spliced two corpora outright because the mixer writes in million-token round-robin blocks. Invisible to every metric being optimised — visible only by reading the mixer. Fixed with a block-diagonal document mask behind a config flag, pre-registered as an experiment at the same 0.010-nat bar that once rejected a loss-saving change.

  8. A long-form probe, and a source that cannot be loaded

    The LAMBADA hypothesis gets measured properly: what share of each source’s documents is at least as long as the context. The first source tried, a famous book corpus, turns out to ship a loading script that modern tooling refuses to execute, and the two obvious parquet mirrors of it declare no licence at all. Only one candidate passes both tests, at 99.3% long documents — but its median document is 104,719 tokens, one hundred times the context, so the passing number is also a warning.

    The training mixture as a stacked bar, beside the share of each source whose documents reach the 1,024-token context: 26.2% for FineWeb-Edu, 0.03% for TinyStories, 13.2% for Cosmopedia, and 17.7% for the mixture overall.
    What that probe found, and it is the reason LAMBADA moved: a fifth of the corpus supplies 0.03% of its long documents.
  9. Session three, and the run finishes

    The final session runs 6.35 hours to complete all 4,770 steps — 2.50 billion tokens, 477 contiguous log rows, no gaps. Validation loss reaches 2.8248 against 3.0533 for the previous version on the identical 100-batch protocol: 0.2285 nats better, 4.6 times the pre-registered bar. But the curve has been flat since step 3,250, so the last 31% of training bought nothing, exactly as the first version’s curve did.

Day 5 2026-09-29

  1. The gate never opened

    The architecture arm finished 0.0436 nats ahead of control — 4.4 times its bar, the largest effect any architecture arm has produced here — and nobody could say why. Measured on a hundred real validation batches, the learned gate sits at 0.0192 against a 0.018 floor and a 0.05 threshold. It never opened. That conclusion came from a single average, which is not a safe inference, so it was pre-registered as a test aimed at my own reading before any GPU time was spent; variation in response to input turned out to be 5e-06. The mechanism is not inert on average, it is inert everywhere.

  2. Our accuracy bars were below our own noise

    The same arm passed its accuracy prediction — 1.2 points on WinoGrande against a half-point bar — and the pass was worthless. The standard errors on these benchmarks at 500 samples are about two points, so the bar sat four and a half times under the noise and the result was 0.54 sigma: the expected size of nothing. The same flaw was in the promotion rule for the version that had just finished, which allowed one point of regression. The thresholds stay exactly as registered and the noise floor is published beside the verdict rather than used to rescue it. But the interim regressions now split cleanly: ARC-Easy and LAMBADA are real, and the other three were inside the noise all along.

  3. Reading the licence changed the answer

    I had assumed the sharing licence on the children’s stories blocked redistribution, and said publish the recipe rather than the data. Reading it, that was wrong. It explicitly places results — the outputs of training, our weights — beyond any obligation, and grants the right to train outright. So all three text sources permit both training and weight release, and this version’s corpus is shippable under three cheap attribution conditions. The asymmetry is unflattering: the licence-clean checkpoint is the one that scores worse, and the earlier version, with copyleft in its training data, is the one whose bytes will never be published.

  4. The missing fifth task

    The registered evaluation came back with four of five benchmarks. LAMBADA had silently failed to load: the dataset still ships a loading script and modern tooling refuses to run those. The data was never missing — six parquet files, the standard 5,153 examples — only the loader path the harness requested was dead. Fixing it took five attempts and four different wrong turns, each an assumption I had not checked, and the fix now lives in the repository with tests rather than inside a notebook. Final answer: 20.8, a 2.2 point drop.

    Grouped bars for five zero-shot tasks, both versions, each with a grey band showing plus or minus one standard error of about two points. Three movements fall inside their own band.
    The completed evaluation. The grey band is the measurement’s own uncertainty, which is why three of the five movements cannot be called a change.
  5. Four kernels died behind green uploads

    The dataset publisher defaults to skipping directories, so the code dataset had been shipping with no config folder, no eval folder, and no JSON at all — which is why a measurement had never been reachable even though the code supported it. Every upload reported success. Then four kernels in a row failed on unchecked assumptions: a missing file, an import used before it was defined, a directory that was never there, and a parser handed markdown where it expected JSON. The publisher now stages one layout, always zips it, refuses an incomplete payload, and then downloads the result back to check seventeen files really are there. Assert the interface, then trust it.

Appendix B. Every decision

41 numbered entries, including the rejected ones

D1Kaggle as the platform: 30 GPU-hours per week, 2x T4held
D2From-scratch only; no fine-tuning, no distillationheld
D358M params: 10 layers x 512 dim, context 1024held
D4TinyStories + FineWeb-Edu corpusheld
D5One change per version, incremental ladderheld
D6Micro-ablation before every full runheld
D7Pre-register experiments with hashed predictionsheld
D8Journal every session; three sources per research claimheld
D9Micro-batch 8 x accum 32: OOM fix at equal effective batchheld
D10Browser-started runs only: API runs cannot read secretsheld
D11Freeze v0.1.0 as the equal-token baselineheld
D12Screen changes at 30M params, 50M tokens, ~30 minheld
D13Journal everything; verify every claimheld
D14Implement Muon+ as experiment E1validated, below threshold
D15E1 at -0.015 misses the 0.02 threshold; do not promoteenforced
D16Fix resume race: single downloader, atomic swap, barrierfixed
D17Expert reprioritization: sweeps first, recipe lock, data before architectureheld
D18Finish v0.1.2 instead of abandoning at 71%held
D19v2 mixture: 60/20/15/5 FineWeb-Edu / TinyStories / Cosmopedia / codebuilt
D20Punch above weight: token efficiency and data quality before scaleheld
D21Distillation rejected: KALIA stays a pure from-scratch lineageclosed
D22Adopt test-time training and self-teaching research trackqueued
D23Invention target: entity memory + in-place test-time weightsqueued
D24Looped depth promoted to front of queuetested, rejected
D25RL self-play and MoE/DSA rejected at this scale, on evidenceclosed
D26Continual-learning recipe planned for post-trainingplanned
D27Publish minimum footprint, honest only, never redistribute dataenforced
D28Rebuild v2 code shard with a per-file license filterdone
D29Build an evaluation harness: probes, bpB, comparison tooldone
D30Reasoning policy: public methods, our own weights, no RLVR yetheld
D31Freeze Muon LR at 0.02; sweep completeclosed
D32Three ancient-text-inspired designs queued (Kautilya, Utsarga, Apoha)queued
D33Four more queued; Jata-patha becomes the reversal experiment X16queued
D34Three Veda-derived designs recordedqueued
D35Gita compilation adapted; prior art cited, not claimed as novelqueued
D36Entity-consistency harness built; novelty claim narrowed after prior-art checkactive
D37Chakravyuha dismissal corrected; Abhimanyu-gap metric builtactive
D38Hash-anchored pre-registration and pre-registered stopping adoptedactive
D39Strategic audit: finish, test, publish, then v0.2.0; freeze the queueheld
D40Quota-exhaustion response: defer the experiment, publish assets nowsuperseded
D41Stop v0.1.2 at step 3,478: converged within noise, quota to v0.2.0done
D42X16 rejected: reversal made the gap worse on both seeds; frozen recipe to v0.2.0enforced
D43The rebuilt corpus redefines the yardstick: v2b val is canonical, v0.1.2 re-baselined on itactive
D44Licence finding disclosed; publish the filtered corpus plus the filter, never the raw corpusactive
D45Accuracy thresholds must sit above the benchmark noise floor; X18 and S-A were both under-poweredactive
D48No architecture arm from a single seed; the gated-residual line is closed as a null resultactive

Appendix C. Every incident

12 published in full

I1API runs cannot read secrets

Secrets attach per notebook in the UI only. Real runs now start from the browser, with a fail-fast connection test.

I2Config missing in the first GPU run

The dataset uploader skips subdirectories. Configs were flattened into the dataset root and the notebook now asserts they exist.

I3Out of memory at step one

Batch 32 at context 1024 exceeded a T4. Fixed with micro-batch 8 x accum 32 and expandable memory segments.

I4A failed run marked complete

Shell-style commands do not fail notebook cells. Replaced with subprocess calls that assert their exit codes.

I5Repository hygiene

Early drafts carried environment-specific wording and inconsistent authorship. Documentation was normalized.

I6Resume crash (EOFError)

Both distributed workers wrote and read the same checkpoint file. Fixed with a single downloader, an atomic swap, and a barrier.

I7Latent crash on older checkpoints

The optimizer loader dropped keys added after those checkpoints were saved. Defaults are restored, with a regression test.

I8Log history lost across sessions

Each session starts with an empty workspace, so pushing logs replaced them. Resume now pulls and appends. The first session rows are recorded as a known loss.

I9Mix step failed after two hours

Relative paths resolved against the wrong directory after a cd. Absolute paths now; a mix-only kernel reuses the completed shards.

I10Wrong mount path and stale code

Dataset mounts live under /kaggle/input/datasets/<owner>/<slug>/, and the dataset had not been re-versioned. Both corrected.

I11A misreported duration

I described a 44-minute ablation as having run five hours, trusting my sense of time over the run-start timestamp. The correction is in the journal.

I12Weekly GPU quota exhausted

The push failed with 30 of 30 GPU-hours used; the UI earlier read "24 hours remaining". Only the push attempt is authoritative. Experiments deferred to the weekly reset.

I13A queue, not a bug

Two kernels sat QUEUED for 45 minutes at Sunday peak capacity, and a retry loop pushed the account past its limit — the cap of two batch GPU sessions counts queued ones, so a retry that assumes a failure can consume the retry's own budget. Deleting the duplicate and pushing once started the run immediately.

I14The published loss curve was missing its first half

The v0.1.2 training log on Hugging Face began at step 1740: a code change adding log-restore-on-resume landed at 06:00 UTC, one session after that session had already started with the old code, so 347 logged steps vanished from the public record. Rebuilt from the repository's own commit history, verified contiguous (steps 10 to 3470, no gaps) and republished. The recovered curve supports the stop decision: the descent finishes by step 1,500 and the remaining 1,750 steps oscillate inside noise.

I15Training windows ignored document boundaries

The corpus is EOS-delimited — all four sources, median document ~200 tokens — but windows were sampled with no document awareness and attention ran unmasked, so nearly every 1,024-token window crossed a boundary. Worse, the mixer writes the four corpora in million-token round-robin blocks, so 0.3% of windows splice two corpora with no separator at all. Invisible to loss and to every benchmark we tracked: found by reading the mixer. Fixed with a block-diagonal document mask behind a config flag, and pre-registered as an experiment (X17) at the same 0.010-nat bar that once rejected Muon+.

I1641.7% of the code corpus is copyleft, and it shipped

A 20,000-file sample of the code corpus measures 41.7% of characters under non-permissive licences, with GPL-family terms covering 39.5% of all files. The licence filter was written, tested, and correct — it just landed 40 minutes after the corpus that needed it, and nothing tied a new filter to a rebuild of the datasets already built. v0.1.2 therefore trained on roughly 50M copyleft tokens, about 2% of its training set, and its published card had listed three of the four sources while omitting the code slice entirely. Eleven minutes of CPU and one random sample found what no loss curve could. The v0.1.2 card now discloses it, and publishing training data is redefined as publishing the filtered corpus plus the filter — never the raw corpus, because redistributing copyleft text is the step that actually triggers the obligation.

I17The dataset was silently dropping two directories

kaggle datasets version defaults to --dir-mode skip, which ignores subdirectories entirely. The code dataset had been published flat, so configs/ and eval/ were both absent and it contained zero JSON files — meaning no kernel could resolve eval/probe_sentences.json, which is the real reason the Abhimanyu gap was never measurable even though the code supported it. The upload reported success every time. Four kernels failed in a row behind a green upload, each from one unchecked assumption: a missing file, an import used before it was defined, a missing directory, and a parser handed markdown where it expected JSON. The dataset publisher now stages one layout, always uses -r zip, refuses to publish an incomplete payload, then downloads the result back and asserts seventeen required files resolve. Assert the interface, then trust it.

I18A lint gate that had never run

The notebook check ran pyflakes via subprocess and printed its output or the word clean. pyflakes was not installed, so the error went to stderr, stdout was empty, and the fallback printed clean — an absent checker reporting success, recorded in several commit messages as a result. It cost a kernel run to expose. The replacement probes for the checker first and refuses to emit a verdict without it, and both of its own failure paths are tested by deliberately breaking them.

KALIA is a personal research project. Not affiliated with any government scheme, company, or other project using a similar name. Training data is never redistributed; every source is attributed in the model card.

Case study · Technical write-up · Back to the portfolio