There Is No Autograd Here
Paired layer functions, mean-loss seeds, AdamW on a flat parameter tape, and the 41-step sanity loop I read but have not yet run.
Index / term
Paired layer functions, mean-loss seeds, AdamW on a flat parameter tape, and the 41-step sanity loop I read but have not yet run.
train_gpt2.py as nanoGPT-shaped model, token-river loader, and a write_state bridge that dumps weights, grads, logits, and loss.