Sep 30th ~ Oct 6th
#128 Latest AI Research Explained Simply
🗞️ Industry News in 1 Line
♥ 6.7k Reflection AI has launched Beam, which is a 501B-parameter MoE model with 23B active parameters. It was pretrained from scratch on 24 trillion tokens and tuned via large-scale reinforcement learning, the open-weights model is slated for release under an Apache 2.0 license alongside FP8 and NVFP4 quantizations. You can apply for early access through their developer platform.

♥ 9.7k Aleph Alpha has released Kolibri, an open-weights MoE model with 78B total parameters, 3.46B active parameters, and support for up to a 1-million-token context window. You can review the technical report and try it on Hugging Face.
♥ 2.8K Google DeepMind has announced Gemini 4 Argon, which is a new frontier model optimized for complex reasoning tasks for coding, cybersecurity defense, and enterprise workflows. It comes with a 1-million-token output limit for extended, multi-step problem solving, the model is currently rolling out to trusted testers via the Fairwind Program ahead of a broader release.

♥ 1.4k Mistral AI has introduced Mistral Large 4 ("Le Chonk"), which is a natively multimodal MoE model with 1 trillion total parameters and 49 billion active parameters. It is built for specialized enterprise workloads such as visual grounding, cyber defense, and finance. The model is available now via the Mistral Cloud API ahead of an open-weights release scheduled for later this month.

The Farmer Who Traded Herbicides For Robots
After his father developed Parkinson’s, founder Clint Brauer began questioning agriculture’s dependence on herbicides. That question led to BOTONY, a robotic farming system designed to enable a more regenerative approach to farming.
Now, investors can join in Greenfield’s next phase of growth through its live Reg A+ offering.
This Reg A+ offering is made available through StartEngine Primary, LLC, member FINRA/SIPC. Please read the Offering Circular and related disclosures before investing. This investment is speculative, illiquid, and involves a high degree of risk, including the possible loss of your entire investment.
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
Sieberling et al. [Massachusetts Institute of Technology, MIT-IBM Computing Research Lab]
♥ 829 Attention
LLMs can process text rapidly at a constant speed, but their memory capacity limits how much context they can recall over long sequences. Expanding this memory causes an explosion in overall parameters which makes models struggle on recall-heavy tasks.

Triadic Gated DeltaNet with a growing second key dimension E at 1.3B parameters.
To overcome this bottleneck, researchers developed triadic linear attention. This approach expands memory capacity in a parameter-efficient way. Modern linear attention networks store information in a two-dimensional matrix formed by combining an input-dependent key vector and a value vector.
The researchers realized they could lift this framework into three dimensions by introducing a second key vector. By storing the product of two distinct keys and a single value into a third-order tensor state, and reading it back out using two matching queries, the architecture multiplies its memory volume by the dimension of that second key while requiring only two modest extra projections.

Forward and backward time of one block of the 1.3B models per sequence, at batch size 4 on H100 GPUs. Triadic GDN with E=8 adds moderate overhead over GDN and is faster than the Transformer beyond 4k tokens.
This 3D memory is fully compatible with high-performance training techniques. The architecture naturally accommodates data-dependent forgetting by applying individual gates to slices of the tensor, and it cleanly removes old associations before writing new ones.
Looped Diffusion Transformer
Chng et al. [SenseTime Research, LeapLab, Tsinghua University, Nanyang Tehnological Univesity]
♥ 551 Looped Transformer bycloud’s pick
Latest image generation models generate high quality images by increasing their size to billions of parameters, which creates computational and deployment bottlenecks.

Looping is a parameter-efficient way to scale text-to-image generation models.
Researchers discovered that instead of stacking endless unique layers, a model can achieve greater computational depth by recycling its existing architecture. They developed the Looped Diffusion Transformer, or Looped-DiT, which routes visual data through a shared group of middle network blocks multiple times within each image-generation step.
Simply repeating layers causes models to blur spatial details and overwrite key information. The team made two innovations: deep supervision, which evaluates and guides predictions across every intermediate cycle during training, and self-modulating attention, which dynamically regulates update strengths so tokens retain critical local details.

Qualitative results of Looped-DiT B/16.
This approach creates a form of silent visual reasoning. As input passes through deeper loops, the network iteratively refines its work. It rearranges objects to satisfy tricky spatial relationships, resolves complex constraints, and self-corrects earlier rendering mistakes without needing explicit text instructions.
In tests, a compact 260-million-parameter Looped-DiT outperformed models roughly six and a half times larger across multiple benchmarks while requiring nearly five times less computation during inference.
Decoding Looped Transformers Better for (Almost) Free
Liu et al. [Apple]
♥ 601 Looped Transformer
Looped Transformers save memory by recycling a single shared block across repeated processing cycles to refine their thoughts. The conventional approach wastes the model's computational power and misses an easy opportunity to catch early errors.

LoopCD lets an earlier loop guide the final one: more accuracy at full depth, less compute at half.
Researchers realized that these discarded intermediate states provide a built-in pair of rough drafts and mature predictions without requiring external training or helper models.
To take advantage of this hidden signal, they introduced LoopCD, which is a technique that contrasts the model's final prediction against an earlier, weaker loop.
Simply steering token selection away from the initial, less-developed guess sharpens its ultimate decision. The method can even adjust this contrast dynamically and applies a stronger push whenever the model hesitates between closely matched candidates to cleanly break ties on tough decisions.

Looped Transformers differ in which part of the network loops.
This adjustment gives a lot of improvements across demanding reasoning and programming tasks, which significantly increases benchmark pass rates on complex mathematics and code generation.

LoopCD applies the contrast before or after the coda and language modeling head.
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
Huang et al. [Institute of Foundation Models, USC, CMU]
♥ 1.2k Looped Transformer
If we can build compact language models by running data through a shared block of layers in repetitive cycles, then it can give us enormous hardware efficiency. However, recurrent architectures have severe computational debt across pretraining, text generation, and reinforcement learning.

Heterogeneous token convergence in Huginn.
Researchers discovered that if a looped model is guided toward a stable resting state, which is known as a fixed point where extra iterations no longer alter internal representations, then the path taken to get there stops mattering.
Once states settle into this equilibrium, the model can safely swap the entire computational history for just the final outcome. This insight unlocks major engineering shortcuts: text generation can share final memory caches across steps to shrink memory footprints threefold without sacrificing accuracy.

The researchers were able to train a distilled companion model which could process initial prompts nearly 1.8 times faster, and use RL updates to cut evaluation runtimes in half by reusing saved endpoints instead of recalculating full loops.
To reliably form these well-behaved endpoints, the researchers addressed two training flaws. First, because different words settle at different speeds, the team created a dynamic depth schedule that learns from prediction feedback, to ensure the model sees properly stabilized context while maintaining enough variation to generalize.
Second, they introduced an orthogonal input method that injects data cleanly at each step, which prevents new inputs from unintentionally canceling out or amplifying existing representations.

