The Frontier

New models, papers, and benchmarks

AI beats Stratego's greatest player on an $8,000 budget

The Frontier

AI beats Stratego's greatest player on an $8,000 budget

One paper closes the last big board-game frontier for AI — at a university price tag — while two disclosures show how the buildout is actually wired: agents leaking what they see, and a chip supplier financing its own customer. Ataraxos, a Stratego AI built by Carnegie Mellon, NYU, Stanford and MIT researchers, beat Pim Niemeijer — the most decorated player in the game's history — 15 wins, one loss and four draws across a 20-game series, the first superhuman result in the game's history, accord

One hour in a scanner is now enough to rebuild what you see

The Frontier

One hour in a scanner is now enough to rebuild what you see

Today's inbox pairs two sides of the same question: what AI can infer about people, and what it takes to defend against it. A brain decoder that used to demand 40 hours in a scanner now needs one, and the founder of Mandiant banked a monster round for AI that attacks like a human team. A brain decoder can now rebuild what you're looking at from a single hour of fMRI — and it learned mostly from images no one ever showed a subject. Michal Irani's team at the Weizmann Institute of Science split t

Google's Gemini 4 Argon gets a 1M-token output limit

The Frontier

Google's Gemini 4 Argon gets a 1M-token output limit

Gemini 4 Argon's output ceiling jumps to one million tokens. Google confirmed in its launch post that Argon can now generate up to one million output tokens, up from the 64K limit on prior Gemini models — an industry first at that scale. Pricing lands at an introductory $2 per million input tokens and $10 per million output, with cached input 95% off, rising later to $4 and $20. The catch: Argon isn't broadly available yet — Google is rolling it out first to a small group of trusted cyber defend

Ten Claude agents formalize a 122-year-old physics problem

The Frontier

Ten Claude agents formalize a 122-year-old physics problem

A proof no human wrote, a 320-billion-parameter release from an unknown Chinese lab, and OpenAI's billing system quietly starting to carry other labs' models — the overnight lane had a research flavor. Ten Claude Sonnet 5.5 agents spent roughly 15 hours and 1,270 messages producing a 17,895-line Lean proof that settles the N=7 case of the Thomson problem — J.J. Thomson's 1904 question of how seven charged points arrange themselves on a sphere to minimize their mutual repulsion. Vals AI, the ben

Google employees question Gemini 4 Argon's real-world coding

The Frontier

Google employees question Gemini 4 Argon's real-world coding

Google shipped its return to the frontier today — and its own people are already arguing about how far ahead it actually is. Plus a public spying accusation between two coding-agent rivals, and another mega-round for AI in hardware design. Some Google employees say Gemini 4 Argon tops the benchmarks but struggles with real-world coding — and Google is disputing that characterization. Bloomberg reported Wednesday that employees with direct access to the model describe its coding abilities as une

AI 101 — What is overfitting?

The Frontier

AI 101 — What is overfitting?

Overfitting is when a model learns its training data too well — it gets the examples it has already seen almost perfect, then performs worse on anything new. Since the entire job of a model is to handle cases nobody showed it, overfitting is the most common way an AI system looks brilliant in a demo and disappoints in production. The term shows up in almost every model release, fine-tuning guide and benchmark write-up, and it is rarely explained. The short version: a model is only worth anythin

Kyunghyun Cho launches Ortet, a frontier AI health lab, with $500M

The Frontier

Kyunghyun Cho launches Ortet, a frontier AI health lab, with $500M

A co-author of the attention mechanism spent five years building AI for drug discovery inside Genentech. Today he launched his own lab — with a half-billion dollars behind it, one disclosed backer, and no priced round. Ortet launched on September 29 as a frontier AI lab for health, with a $500 million commitment from Thoreau and a founding team led by Kyunghyun Cho. Cho co-developed the attention mechanism and the gated recurrent unit — the groundwork under every transformer in production today

AI 101 — What is reinforcement learning?

The Frontier

AI 101 — What is reinforcement learning?

Reinforcement learning is how a machine learns by trial and error: it tries things, notices which attempts earned a reward, and drifts toward the behaviour that earns more of it. There is no answer key involved. In the supervised learning most people picture, someone hands the model millions of labelled examples and it learns to copy them. In reinforcement learning, nobody says what the right move is — the machine acts, gets a number back telling it how well that went, and adjusts. The classic

Anthropic ships Sonnet 5.5 as cyber safeguards move down a tier

The Frontier

Anthropic ships Sonnet 5.5 as cyber safeguards move down a tier

Anthropic's second Claude 5.5 model lands with frontier-class benchmark deltas at a mid-tier price — and the safety machinery that used to sit only on the top model. Meanwhile a chipmaker buys into the financing layer of the AI buildout, and Congress gets its most detailed AI-safety bill yet, with no path to a vote. Anthropic released Claude Sonnet 5.5 on September 28, and the interesting part is not the score — it is that Sonnet 5.5 beats Anthropic's own top model on the hardest agentic evalua

The Take — A refusal rate is a capability score

The Frontier

The Take — A refusal rate is a capability score

Artificial Analysis put out a serious benchmark this week and buried its most interesting result in a second chart. I think that is the wrong way round. A model that declines 38 percent of a job does not have less capability on the other 62 — it has a policy. The Cyber Index's headline score turns that policy into a capability gap, silently, and the headline is the thing that spreads. Here is the mechanism, because this is an arithmetic complaint, not a philosophical one. The index is an equal-

H company opens Holo4: a 27B agent that drives your desktop

The Frontier

H company opens Holo4: a 27B agent that drives your desktop

Two releases in the same morning frame the agent debate neatly: a small open-weight model that can take over a desktop cheaply, and a safety stack being sold on silicon because the model-level guardrails keep failing. H company released Holo4, a computer-use agent family in two open-weight sizes — 27B dense and 35B-A3B mixture-of-experts — alongside Holo4's weights in BF16, FP8, NVFP4 and 4-bit GGUF formats on Hugging Face. The pitch is interface-agnostic: the same model clicks and types on a s

LLM4MIP says it closed 34 open optimisation problems — MIPLIB lists 29 as incumbents

The Frontier

LLM4MIP says it closed 34 open optimisation problems — MIPLIB lists 29 as incumbents

A research site says an LLM workflow closed 34 open optimisation problems. The benchmark that would have to agree lists 29 entries from the same team — and none of them as proven optimal. The gap is the story. An LLM-assisted workflow now claims 34 of 132 open MIPLIB instances "resolved" — but MIPLIB's own log carries 29 of those submissions, all as incumbent improvements, not proofs. The work is LLM4MIP, published as a research website on 21 September by Yicheng Huang and Wenzhi Gao as co-lead