Sep 22nd ~ Sep 29th
#127 Latest AI Research Explained Simply

🗞️ Industry News in 1 Line

  1. ♥ 53k Anthropic has released Claude Sonnet 5.5, which comes with improved benchmark performance, enhanced cyber safeguards, and faster generation speeds. The updated model is optimized for routine coding, UI design, and structured document creation while operating with greater token efficiency to reduce per-task costs. You can test the model via the Anthropic API.

  2. ♥ 17k ElevenLabs has released Eleven v4 and Eleven v4 Turbo. It comes with acoustic control via inline tags for emotion, pacing, and sound effects across more than 90 languages. The update also adds International Phonetic Alphabet (IPA) support for exact pronunciation tuning, alongside enhanced dynamics for multi-speaker conversational agents. You can test the models via the ElevenLabs API or explore integration examples and SDKs on GitHub.

Intuitive AI Academy - NEW Quiz & Chapters!

My project Intuitive AI Academy just had an upgrade!

We’ve added a new quiz system that tests your understanding after every chapter, along with a weekly leaderboard to keep you accountable and make learning a little more competitive.

We also just released a new chapter on Linear Attention, covering its history, the key techniques behind it, and how it’s being used in frontier models today.

We currently have a back to school offer, where you would get 50% off on the yearly plan.

Use code: BACKTOSCHOOL

Self-Play Pretraining with Zero Data

Cowsik et al. [Tel Aviv University, Stanford University, LAPTh]

♥ 3.3k Pre-training

AI models rely on human-curated data and this creates a bottleneck where progress is ultimately limited by the bounds of human knowledge rather than computing power. If machines can instead learn to invent their own training material from scratch, intelligence could scale without human limits.

This paper introduced a framework where two randomly initialized models teach each other without using a single byte of human text or media. One model acts as a generator, writing tiny computer programs for a universal virtual machine that executes them into raw byte streams.

Self-Play Pretraining with Zero Data

A second model predicts these bytes step-by-step. To ensure the generated data is genuinely useful rather than random noise, the generator receives a reward whenever it crafts programs right at the frontier of what the learner is currently capable of understanding. This creates an adaptive, self-guided curriculum that constantly moves forward as the learner improves.

Even though neither model ever viewed natural data during training, the learner showed predictable, power-law improvements when tested on completely unseen human text, computer code, speech, and images.

Memory Attention

Kang [Independent]

♥ 598 LLM Attention

AI models need massive computing power because their internal attention mechanisms constantly recalculate information, even when many word representations could simply be reused across contexts.

Standard Attention, Memory Attention, and MA-Offload.

To solve this, researchers introduced Memory Attention, an architecture that removes the traditional, computationally dense value calculations. Instead of this, the model constructs its values by combining dynamic contextual keys with layer-specific token memory tables.

Here, each layer maintains a reference table of learned word representations. Instead of spending compute to generate these values, the model simply looks up the token in its stored memory and adds it directly to the contextual key. During inference, preparatory normalization math can be precomputed and folded right into the tables, which can streamline the entire value-creation pipeline down to a fast memory lookup and basic addition.

Training and inference comparison of MA and Standard

In this design, the memory lookups depend only on token identities, the tables do not need to stay locked inside expensive GPU memory. Through an offloading strategy, the memory tables can reside in standard CPU memory and be smoothly prefetched to the GPU just moments before each layer needs them.

This technique allowed models to more than double their total parameter capacity while actually requiring less GPU parameter storage than standard architectures.

JEPA-Anything: Learning Predictive Models across Different Worlds

Cui et al. [PhAI Labs, The Chinese University of Hong Kong, Fudan University, City University of Hong Kong, University of Bristol, Stanford University, University of Oxford, Princeton University]

♥ 696 JEPA

AI world models have to be custom-built from scratch for every specific domain, from molecular simulations to medical forecasting. Traditional architectures can cram an entire simulated future into a single snapshot which can cause obvious patterns to drown out subtle information.

To solve this, researchers introduced JEPA-Anything, a universal world-modeling framework powered by a mechanism called orthogonal predictive factorization. Rather than forcing one predictive pathway to forecast every detail at once, the system breaks the target state into separate, complementary slices.

In this dedicated pathways predict each slice independently, which prevents different physical effects from competing with one another, before stitching them back together into a cohesive whole. By using simple adapters to translate raw inputs into a common format, the same predictive engine can analyze completely different environments without changing its core design.

Scenario atlas for JEPA-Anything

The researchers tested this method across seven distinct domains, including physical fields, weather, control systems, and clinical records tracking over a thousand disease events. Compared to matched baselines, JEPA-Anything improved performance across all ten tested dynamics benchmarks, reduced intervention prediction error on a digital testbed by nearly 35%, and recorded the lowest errors across short- and long-horizon molecular rollouts.

Factor-wise functional interventions.

Contrastive World Models

Li [Independent]

♥ 1.8K World Models

AI should be able to anticipate the future and plan their actions effectively. However, traditional world models understand their surroundings by predicting and reconstructing every single pixel on screen. This strategy falls apart when moving distractions or cluttered backgrounds consume an agent's focus with irrelevant details.

Pixel-prediction-based World Models (left) and Contrastive World Models (right).

To solve this, researchers introduced Contrastive World Models, a new approach that removes the pixel-generating component entirely. It uses the Dreamer framework and this method replaces full visual reconstruction with an objective that maximizes the mutual information between an agent's state-action history and local feature patches of future observations.

Rather than attempting to draw an exact replica of what comes next, the system learns representations that preserve only the information predictive of future events. This encourages the model to track the dynamics essential for planning while ignoring visual noise that has no bearing on the task.

Architecture of Recurrent State Space Model

Reply

Avatar

or to participate