Feb 8, 2026I Pretrained a 360M LLaMA-Style Language Model from Scratch on 6B FineWeb Tokens (Single H100)Kashif
Nov 18, 2025From Karpathy's micrograd to smoltorch: Understanding Autograd from First PrinciplesKashif