Logo
Search
The AI Timeline
LOG IN
HOME
ARCHIVE
TAGS
AUTHORS
UPGRADE
Logo
by cloud
Explorative Modeling: Third Pre-training Axis?

/

Aug 4, 2026

Explorative Modeling: Third Pre-training Axis?

Plus more about Memory Foundation Model, Weak-to-Strong OPD, and LeRoPE

by cloud
by cloud
Kimi K3 Technical Report

/

Jul 29, 2026

Kimi K3 Technical Report

plus more about Hilbert Operator for Progressive Encoding, RLVR-Native Optimization Stack, Soap & Muon At Scale, and Measuring Reward-Seeking

by cloud
by cloud
On-Policy Delta Distillation

/

Jul 21, 2026

On-Policy Delta Distillation

plus more about Concurrent Image Understanding and Generation, Latent and Explicit Reasoning with Looped Transformers and more

by cloud
by cloud
Why Memorized Knowledge Fails to Generalize in LLM Finetuning

/

Jul 14, 2026

Why Memorized Knowledge Fails to Generalize in LLM Finetuning

Plus more about Single Async Opt for Agentic RL, Remember When It Matters, and Sparse Delta Memory

by cloud
by cloud
You Only Need 1 Layer for RLVR?

/

Jul 7, 2026

You Only Need 1 Layer for RLVR?

plus more about AdaJEPA, Program-as-Weights, The World Is In Your Mind, and Dual On-policy Distillation

by cloud
by cloud
DeepSeek Just dropped a new speculative decoding method!

/

Jun 30, 2026

DeepSeek Just dropped a new speculative decoding method!

plus more about Tapered LMs, Improved LLDMs, AutoData, and You Don't Need To Run Every Eval

by cloud
by cloud
What even is a >< former (yes >< former)

/

Jun 23, 2026

What even is a >< former (yes >< former)

plus more about Looped World Models, Fixed-Point Reasoners, and ExpRL

by cloud
by cloud
MiniMax M3's New Attention: MiniMax Sparse Attention

/

Jun 16, 2026

MiniMax M3's New Attention: MiniMax Sparse Attention

plus more about FlashMemory-DeepSeek-V4, Trajectory-Refined Distillation, Test-Time Gradient Guidance, and End-to-End Context Compression at Scale

by cloud
by cloud
Microsoft just shared the frontier data engineering secrets

/

Jun 9, 2026

Microsoft just shared the frontier data engineering secrets

plus more about If LLMs Have Human-Like Attributes, Then So Does Age of Empires II, Cosmos 3, and Robots Need More than VLA and World Models

by cloud
by cloud
DiffusionBlocks: Save 2-3x Training Memory!?

/

Jun 2, 2026

DiffusionBlocks: Save 2-3x Training Memory!?

plus more about Bitter Lesson in Data Filtering, Do Language Models Need Sleep, and Neural Weight Norm.

by cloud
by cloud
Generative Recursive Reasoning

/

May 26, 2026

Generative Recursive Reasoning

plus more on the Benefits of Subword Tokenization, HRM-Text, Probabilistic Tiny Recursive Model, and Vector Policy Optimization

by cloud
by cloud
Long Context Pre-Training w/ Lighthouse Attention

/

May 19, 2026

Long Context Pre-Training w/ Lighthouse Attention

plus more about Self-distilled Agentic RL, Embedded Language Flows, and Negation Neglect

by cloud
by cloud
Think In Diffusion: Continuous Latent Diffusion Language Model

/

May 12, 2026

Think In Diffusion: Continuous Latent Diffusion Language Model

plus more on Sparser, Faster, Lighter Transformer LMs, Manifold Steering, and Teaching Claude Why

by cloud
by cloud
DeepSeek's Deleted Paper: Thinking With Visual Primitives

/

May 5, 2026

DeepSeek's Deleted Paper: Thinking With Visual Primitives

can't believe they removed this paper unknowningly

by cloud
by cloud
Kimi Moonshot: Prefill-as-a-Service!?

/

Apr 21, 2026

Kimi Moonshot: Prefill-as-a-Service!?

plus more about Looped Transformers, Nexus, RNN with Memory, and more

by cloud
by cloud
Neural Computer: Running an OS within an AI?!

/

Apr 14, 2026

Neural Computer: Running an OS within an AI?!

plus more about In-Place TTT, TriAttention, and Interleaved Head Attention.

by cloud
by cloud
Embarrassingly Simple Self-Distillation Technique

/

Apr 7, 2026

Embarrassingly Simple Self-Distillation Technique

plus more on Path-Constrained MoE, HISA, and Screening is not enough

by cloud
by cloud
LeWorldModel: JEPA but more practical

weekly papers recap

/

Mar 31, 2026

LeWorldModel: JEPA but more practical

plus more on Claudini, Composer 2, and self-distillation

by cloud
by cloud
Rotate attention by 90 degrees...? Kimi's New Attention Residuals

/

Mar 25, 2026

Rotate attention by 90 degrees...? Kimi's New Attention Residuals

plus more about V-JEPA 2.1, Mamba 3, and latent planning

by cloud
by cloud
You can train OpenClaw just by talking to it?

/

Mar 17, 2026

You can train OpenClaw just by talking to it?

and more about GLM-OCR, pre-pre-training on NCA, IndexCache, and neural thickets

by cloud
by cloud
Flash Attention 4 is nuts

/

Mar 10, 2026

Flash Attention 4 is nuts

and more about Speculative Speculative Decoding, SWE-CI, and Beyond Language Modeling

by cloud
by cloud
Compress Context... Into a LoRA!?

/

Mar 4, 2026

Compress Context... Into a LoRA!?

plus more on Learning Without Training and The Geometry of Noise

by cloud
by cloud
Google Presents A Brand New Way To Train Latents

/

Feb 24, 2026

Google Presents A Brand New Way To Train Latents

plus more about Experiential RL, GLM-5 Report, and Attention Matching

by cloud
by cloud
Using Diffusion To Interpret LLMs?! Generative Latent Prior

/

Feb 17, 2026

Using Diffusion To Interpret LLMs?! Generative Latent Prior

plus more on Evolving Agents via Recursive Skill-Augmented RL and Low Hanging Fruits in Vision Transformers

by cloud
by cloud
New Generative Paradigm: Drifting Model

/

Feb 10, 2026

New Generative Paradigm: Drifting Model

an insane big week in AI reseasrch

by cloud
by cloud
Load more

The AI Timeline

Follow The Latest Cutting Edge AI Research in 5 minutes a week.

© 2026 bycloudai.
beehiivPowered by beehiiv