All Posts
/
May 12, 2026
plus more on Sparser, Faster, Lighter Transformer LMs, Manifold Steering, and Teaching Claude Why
May 5, 2026
can't believe they removed this paper unknowningly
Apr 29, 2026
plus more about Hyperloop Transformer, Qwen-3.5 Omni, and Scaling Self-Play with Self-Guidance
Apr 21, 2026
plus more about Looped Transformers, Nexus, RNN with Memory, and more
Apr 14, 2026
plus more about In-Place TTT, TriAttention, and Interleaved Head Attention.
Apr 7, 2026
plus more on Path-Constrained MoE, HISA, and Screening is not enough
weekly papers recap
Mar 31, 2026
plus more on Claudini, Composer 2, and self-distillation
Mar 25, 2026
plus more about V-JEPA 2.1, Mamba 3, and latent planning
Mar 17, 2026
and more about GLM-OCR, pre-pre-training on NCA, IndexCache, and neural thickets