Frontier agents flunked two NeurIPS papers — graded by the authors

Share
Frontier agents flunked two NeurIPS papers — graded by the authors

Two very different pictures of where agents actually are today: still failing at the top of AI research, and already handling real work inside a car cockpit and ByteDance's enterprise stack.

A new study turned frontier agents loose on two unpublished NeurIPS 2026 submissions, gave each six days and thousands of dollars of compute, then had the papers' original authors grade the output. The agents completed all of the engineering without human help and still could not make substantial progress on the research questions — both papers were rejected outright. The authors' post-mortem names five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses when the research design ran into trouble, ineffective backtracking from dead ends, poor awareness of their own resource budget, and instruction drift. A second model and scaffold reproduced the same failures.

The method is the interesting part. Current evaluations either test narrow, verifiable tasks — which excludes open-ended research by construction — or submit AI-written papers to blind peer review, which the authors call overstretched and stochastic. Their alternative, "shadow evaluations," hands an agent the central question of a real unpublished paper and lets the people who wrote that paper judge the result. They released the expert reviews, survey responses, agent repositories and logs.

Read the conclusion carefully, because it is narrower than the headline circulating around it. The paper says today's agents can do the engineering of AI research but struggle with the research lifecycle. It does not claim autonomous self-improvement is impossible — the models tested were already a generation behind the frontier when the runs happened. We made a related argument earlier this week — An agent grading its own homework is an alibi, not proof — and this study is the same lesson from the other direction: capability shows up first in the parts of research that can be checked cheaply.


BYD and Alibaba Cloud put Qwen at the centre of the car's brain, launching a "super agent" cockpit rather than a better voice assistant. At its September 14 event, held alongside the all-electric Denza N8L launch, BYD showed an architecture where a cloud hub runs Qwen for intent recognition, task decomposition and service dispatch, then hands execution to separate agents for vehicle control, music, video, search and outside services. The cockpit can call Alibaba ecosystem services including Fliggy and Taobao Shangou, and BYD is opening an agent platform so third-party developers and AI agents can plug into the car. The pitch from BYD senior vice president Yang Dongsheng is orchestration, not conversation: the system is meant to finish a task — search, routing, booking — without the driver switching apps, and the two companies have been working together since 2023. Worth watching whether outside developers show up for a platform whose distribution depends on one automaker's fleet.


ByteDance's CEO said the company has merged Doubao, Feishu and Volcano Engine and will pour more resources into the enterprise market. Speaking at the Feishu Future Unlimited conference in Beijing, Liang Rubo argued that once agents become usable, the agent, the model and the collaboration environment have to fuse to be worth anything — which is why the three units were consolidated two months ago. His division of labour: Doubao Work supplies the intelligence, Feishu carries the company's context and tools, and Volcano Engine lets enterprises build their own agents on ByteDance models and compute, then wire them back into Feishu to work with people and other agents. It is a direct shot at Microsoft's Copilot stack, and it makes ByteDance an enterprise software vendor more than a consumer app company.

What to watch: whether shadow evaluations get adopted as a standing, re-run benchmark with each new model release — the authors' own suggestion is that the task, harness and budget stay frozen so only the model changes.

If a lab published an agent that reproduced your paper, would you believe the result before you read the code? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t