All Posts
/
Dec 30, 2025
plus more on Self-Play SWE-RL, Step DeepResearch, and Attention Is Not What You Need
Dec 23, 2025
Next-Embedding Prediction Makes Strong Vision Learners, Let's (not) just put things in Context, Spherical Equivariant Graph Transformers, and moree
Dec 16, 2025
Scaling Up Diffusion Language Models to 100B, Adding 1 Attention Layer & Make Visual Encoders Generate Images, LayerNorm Is Not Needed In Transformer, and more
Premium Insights
Dec 11, 2025
Basically recapping what I missed in the last 4 months
Dec 9, 2025
PretrainZero, Stabilizing RL with LLMs and more
Dec 4, 2025
Breaking down "Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?"
Dec 2, 2025
FreeFlow, DeepSeekMath-V2, Soft Adaptive Policy Optimization, and more
Nov 27, 2025
Plus more on Seer, Virtual Width Networks, SAM 3, and Evolution Strategies at the Hyperscale
Nov 18, 2025
LeJEPA, The Path Not Taken, and more