All Posts
/
Jan 13, 2026
and more on Dead Salmons of AI Interp, GDPO, From Entropy to Epiplexity
Jan 6, 2026
And more about Recursive Language Models, LongCat ZigZag Attention, and LoRA RL
Dec 30, 2025
plus more on Self-Play SWE-RL, Step DeepResearch, and Attention Is Not What You Need
Dec 23, 2025
Next-Embedding Prediction Makes Strong Vision Learners, Let's (not) just put things in Context, Spherical Equivariant Graph Transformers, and moree
Dec 16, 2025
Scaling Up Diffusion Language Models to 100B, Adding 1 Attention Layer & Make Visual Encoders Generate Images, LayerNorm Is Not Needed In Transformer, and more
Premium Insights
Dec 11, 2025
Basically recapping what I missed in the last 4 months
Dec 9, 2025
PretrainZero, Stabilizing RL with LLMs and more
Dec 4, 2025
Breaking down "Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?"
Dec 2, 2025
FreeFlow, DeepSeekMath-V2, Soft Adaptive Policy Optimization, and more