All Posts
/
Feb 18, 2025
Plus more about Continuous Concepts (CoCoMix), and Distillation scaling laws
Feb 11, 2025
Plus more about OmniHuman-1, and Simple test-time scaling
Feb 4, 2025
Plus more about Supervised Fine-Tuning (SFT) vs Reinforcement Learning (RL), and Janus-Pro
Jan 28, 2025
Plus more about Transformer2 and Kimi k1.5
Jan 21, 2025
Plus more about MiniMax-01 and Scaling LLM Test-Time Compute
Jan 14, 2025
Plus more about Towards System 2 Reasoning in LLMs and Memory Layers at Scale
Jan 7, 2025
Plus more about ModernBERT, and Qwen 2.5 Technical Report
Dec 17, 2024
Training Large Language Models to Reason in a Continuous Latent Space, and [MASK] is All You Need
Dec 11, 2024
DeMo: Decoupled Momentum Optimization, and Densing Law of LLMs