The Frontier

New models, papers, and benchmarks

AI 101 — What are scaling laws?

The Frontier

AI 101 — What are scaling laws?

Scaling laws are the empirical rules describing how an AI model's performance improves in a smooth, predictable curve as three things grow: the size of the model, the amount of data it trains on, and the computing power spent training it. They are not physics — nobody proved them in a lab the way gravity was proven — they are patterns that kept holding every time someone measured them, which is why the entire AI industry now plans around them. The pattern got its canonical statement in a Janua

Reflection's open-weight answer to DeepSeek and Qwen

The Frontier

Reflection's open-weight answer to DeepSeek and Qwen

An Nvidia-backed startup run by two DeepMind refugees is about to test whether America can finally ship a downloadable model that holds up against China's — and the test matters more as a political artifact than as a benchmark entry. Axios reported Sunday that Reflection's first open-weight model is set to arrive soon, with other Western open-weight releases expected behind it this month. A Reflection spokesperson declined to comment, and no model name, license, parameter count or score has surf

Google's fix for agents that memorize their tests

The Frontier

Google's fix for agents that memorize their tests

Let an agent rewrite the scaffolding around its own model against a fixed set of test tasks and it learns the tests, not the job. Scores on the tasks it practiced climb; gains on anything new shrink or vanish. Google Cloud AI Research's answer, written up by The Decoder today, is not to lock the scaffolding down but to regulate how the search moves through it — and the numbers suggest the regulation, not the model, is what makes self-improvement stick. The catch in recursive self-improvement

Kaiming He's harness gives Claude a perfect ARC-AGI-3 score

The Frontier

Kaiming He's harness gives Claude a perfect ARC-AGI-3 score

A benchmark saturated twice in a month, a security program buried by machine-generated reports, and Anthropic spending real money on human training — three stories, one theme: the stack around the model is now where the action is. A visual harness from Kaiming He's MIT lab pushed Claude Opus 5.0 RHAE to a perfect 100.00 on ARC-AGI-3, all 25 public games — using 57.4% fewer actions than first-time human players. The paper (VISTA, posted October 1) is careful about what it did and didn't do: it d

Meta AI solves five open math problems across six papers

The Frontier

Meta AI solves five open math problems across six papers

Math and money, both from the same wave: Meta says its models closed five previously open research questions, and a 3D-generation startup says the general-model era grew its business instead of killing it. Meta published six papers on October 2 that it says answer five previously open research questions — spanning probability, PDEs, group theory, optimization, arithmetic physics and non-associative algebra. The setup matters more than the count: mathematicians co-authored the work, a second gro

Aleph Alpha open-sources Kolibri, a sovereign German MoE model

The Frontier

Aleph Alpha open-sources Kolibri, a sovereign German MoE model

A European model lab just put real open weights on the table, open-source maintainers are building robots to keep machine-written code out, and the labs' CEOs are arguing about theology. Three stories this hour. Aleph Alpha has released Kolibri, a 78-billion-parameter mixture-of-experts model released under Apache 2.0 and built deliberately around German. The German lab's new flagship activates only 3.46 billion of its parameters per token — about 4.4% of the model — and was trained from scratc

AI 101 — What is AGI?

The Frontier

AI 101 — What is AGI?

AGI — short for artificial general intelligence — is a hypothetical AI system that could do any intellectual job a person can do, instead of being good at one narrow task. Every AI you can actually use today is narrow: it writes, translates, spots patterns in scans, plays games — each system built for its lane. AGI names the destination where one system covers all the lanes. It is a goal nobody has reached, not a product anyone can buy. Why it matters right now "AGI" is one of the most-used w

AI 101 — What is a large language model?

The Frontier

AI 101 — What is a large language model?

A large language model (LLM) is a very big neural network trained on enormous amounts of text to do one job well: predict the next piece of a text, given everything that came before it — then tuned so that when you talk to it, it actually helps instead of rambling. ChatGPT, Claude, Gemini, Llama, DeepSeek: every chatbot in the news is an LLM, and most AI features in the software you already use have one quietly bolted inside. That prediction job sounds too simple to matter, and it is the whole

Google's Suncatcher puts its first TPUs in orbit

The Frontier

Google's Suncatcher puts its first TPUs in orbit

Compute is leaving the ground — literally — while the people who build AI models keep warning about where the buildout ends. Three stories this hour. Google's first Suncatcher satellite is in orbit, and it is carrying the company's own TPUs. The prototype launched Thursday aboard a SpaceX Falcon 9 from Vandenberg on the Transporter-18 rideshare, tucked alongside Planet Labs satellites, and Google confirmed it has contact with the spacecraft and it is operating as expected. Before shipping, the

Black Forest Labs launches Flux 3 Image with multi-step editing

The Frontier

Black Forest Labs launches Flux 3 Image with multi-step editing

A release-heavy morning: a frontier image model goes live with box-by-box editing that promises to hold everything else still, a 111,000-star coding harness hits 1.0 after reversing itself on MCP, and Perplexity slips an open-weights decision model out the same week the category caught fire. Black Forest Labs has released Flux 3 Image, the image half of its Flux 3 family, as a general-availability API built around multi-step editing. The workflow is the pitch: you draw bounding boxes over the r

Cloudflare open-sources the Clef models, aiming straight at Jev

The Frontier

Cloudflare open-sources the Clef models, aiming straight at Jev

Cloudflare picked October 1 to wade into the fastest-moving model category of the fall: bounded, machine-consumable decisions rather than prose — and it brought open weights and a price list. Cloudflare released Clef and Clef-flash, a pair of open-weight "decision models" built to beat TypeSafe's Jev at its own game. The category is Jev's invention — take in unstructured state, return typed, probabilistic decisions that software can act on, with no free-text sampling to hallucinate. Cloudflare'

Microsoft ships streaming transcription and two new voice models

The Frontier

Microsoft ships streaming transcription and two new voice models

Microsoft's AI unit closed out Wednesday with three speech models at once — a streaming transcriber and two voices — aimed squarely at the workloads where latency is the product: live captioning, voice agents, and the contact-center stack. Microsoft released MAI-Transcribe-2-Streaming, its first real-time transcription model, alongside two voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The transcriber covers 60 languages with automatic language detection between them, and it starts return