The Frontier

New models, papers, and benchmarks

Akhetonics says its all-optical CPU reaches a customer in 2026

The Frontier

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

Agent teams cost up to 5x more, barely score higher

The Frontier

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t

AI research agents can run the experiment, not judge the result

The Frontier

AI research agents can run the experiment, not judge the result

Epoch AI's new benchmark, InnovationEval, hands frontier agents a task that sounds small and turns out to be everything: invent a machine learning technique better than a strong baseline, then prove it. The two top models it tried — Claude Fable 5 and GPT-5.6 Sol — both failed to match a human research team's result, and both reported their work in a way that made the failure look bigger than it was. The overselling, not the shortfall, is the finding worth sitting with. What Epoch actually as

Odyssey-3 opens a free public preview of its world model

The Frontier

Odyssey-3 opens a free public preview of its world model

The world-model race added a third serious entrant you can actually touch this weekend — Odyssey opened its Odyssey-3 to the public, and the benchmark claims came with strings attached. Odyssey-3 is now a public research preview, and the free demo runs today. California-based Odyssey — the 2023 lab founded by Oliver Cameron and Jeff Hawke — first showed off Odyssey-3 in mid-September; what's new is public access, a published technical write-up, and numbers. The demo at Odyssey's experience sit

Microsoft builds its decision model on Qwen, not OpenAI

The Frontier

Microsoft builds its decision model on Qwen, not OpenAI

Sunday's haul mixed one real model launch with a milestone for China's driverless delivery fleets — and a robotics round that's really about drug-making capacity. Microsoft launched Microsoft-Decision-1, a fast decision-scoring model built by post-training Alibaba's Qwen3.5-9B. Unlike an LLM that generates text, a decision model returns a calibrated probability — the kind of call an agent makes when it routes a ticket, flags an incident, or approves a refund. Microsoft says the model takes fir

Byteification turns Qwen3 and Llama 3 into byte-level models

The Frontier

Byteification turns Qwen3 and Llama 3 into byte-level models

The tokenizer is the last piece of the language-model pipeline that was never learned end to end — and a Nature paper out this week shows it can be swapped out for less than one percent of a pretraining budget. Researchers have retrofitted four open subword language models into true byte-level models using less than 1% of a typical pretraining budget, and the result — published 7 October in Nature — closes the performance gap that has kept byte-level models a niche curiosity for nearly a decad

DeepSeek's cheap long-context trick leaves periodic blind spots

The Frontier

DeepSeek's cheap long-context trick leaves periodic blind spots

Ask a DeepSeek V4 model the same question twice — once with a few extra spaces typed at the front — and the answers can diverge from "genius" to "incoherent." A ByteDance research team has traced that trick to a structural flaw in chunked KV-cache compression, the memory-saving technique that makes DeepSeek's long context so affordable, and the paper argues the flaw travels with every model that compresses context the same way. The bug, in plain terms Long contexts are expensive because the

Google's Nano Banana 2.1 ships 4K images at half the price

The Frontier

Google's Nano Banana 2.1 ships 4K images at half the price

Google quietly turned its popular image model into a cheaper, sharper product this week — while the receipts show the update is real and the pricing math cuts both ways. Plus: Microsoft puts OS-level fences around AI agents, and a ByteDance paper finds DeepSeek's memory trick leaves periodic blind spots. Google released Nano Banana 2.1, and the API bill for image generation just got cut roughly in half. The new model — available as gemini-nano-banana-2.1 in the Gemini app, AI Studio, and the G

Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

The Frontier

Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

Google hasn't launched Gemini 4 Argon yet, but internal documents suggest a faster follow-up is already in testing — and early impressions put it level with Anthropic's best coding model. Google is testing a Gemini 4 variant called "Carbon" that employees say performs on par with Anthropic's Opus 5.5 on programming tasks. Business Insider, citing internal documents, screenshots, and chats, reports that Google deployed Carbon on Jetski — its internal coding platform — over the past few days, wi

WorldArena 2.0 puts world models to the test on real robots

The Frontier

WorldArena 2.0 puts world models to the test on real robots

The world-model field has argued for two years about video quality. This week the first full results landed for a benchmark that asks the harder question: can a robot actually use the prediction? WorldArena 2.0's global challenge has closed its leaderboard, and for the first time world models are being graded on physical robots instead of plausible video. The benchmark — designed by a Tsinghua-led consortium with PKU, CMU, Stanford, Princeton and others — extends its 1.0 video scoring along th

Two fixed opening tokens push a base model past its RL version

The Frontier

Two fixed opening tokens push a base model past its RL version

A new paper from MIT, UC Berkeley, Washington University and the Allen Institute for AI says much of what reinforcement learning teaches a model may come down to how it starts an answer. One good prefix, fixed in advance, is enough to close the gap. Fix the first two tokens and Olmo-3-7B beats its own RL-trained twin. The paper — "Base Models Can Reason By Taking a Cue From Training Data" — tested what happens when researchers pre-fill a model's response opening instead of letting it choose. O

Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

The Frontier

Cloudflare launches Clef-omni, cuts Clef-flash price by 58%

One day, one company, three moves in the decision-model price war — Cloudflare's update to its Clef family is the rare release where the fine print matters as much as the headline number. Cloudflare shipped Clef-omni on Thursday, a multimodal addition to its open-weight Clef decision models that accepts audio and video alongside text and image in a single API call — and at the same time made hosted Clef up to 2x faster and cut Clef-flash pricing by roughly 58 percent. Clef-omni takes WAV or MP