Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own.
Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the team scored 7.3 points higher (p=0.005). The starkest number: an Opus 5.5 team at max effort cost $122.17 per app against $4.08 for a single mid-effort agent, roughly 30 times the spend, for a 1.6-point gain that wasn't statistically distinguishable from zero. Most of the waste is structural — every subagent carries its own context, re-sent on every model call, and the median run burned 224 million cached tokens per app. Anthropic's own Opus 5.5 system card shows the same shape from the other direction: teams of 100 reach a given score faster, but after 24 hours they land only slightly ahead of teams of 10, and OpenAI's Noam Brown has said plainly that four agents buy you twice the speed at twice the cost — speed, not quality. The take: parallelism is a latency tool, not an intelligence tool, and anyone running agent swarms for benchmark score is paying the coordination tax for a rounding error.
Aurora, Kodiak, and every other driverless truck maker just lost their biggest paper-work blocker. The Federal Motor Carrier Safety Administration granted a five-year exemption on October 7 letting driverless trucks swap the federally required reflective warning triangles — which a human driver must deploy within 10 minutes of a breakdown — for cab-mounted amber beacons that activate within five minutes. It converts 12 months of stopgap three-month waivers into durable permission, and unlike the earlier waivers it's open to any Level 4 carrier that notifies the agency; Kodiak, Waabi, and Stack AV have already opted in. Aurora took the previous denial to court and appealed to the D.C. Circuit, so the exemption settles the fight it spent a year picking — we covered Aurora's current fleet in Aurora has 20 driverless trucks running and wants 200 by year-end.
Texas is still making data centers wait — and a16z's post-mortem explains why the pause is really about the interconnection queue. The state's grid operator went from 63 gigawatts of large-load requests at the end of 2024 to 474 gigawatts by June, about five times ERCOT's record peak demand and roughly 90 percent data centers, most of them speculative projects that may never plug in a single GPU. Governor Abbott ordered an audit in August; ERCOT then paused Batch Zero approvals for loads of 75 megawatts and above — including 17 projects that had cleared every other step — and a state permit freeze followed, with an audit report and a TCEQ update to the governor due in October. The portfolio angle matters: a16z argues the queue is clogged with duplicate and speculative applications that planners can't tell apart, which is a very different problem from the local opposition politics we wrote about in Texas can't agree on who gets to say no to a data center.
What to watch: the October 19 TCEQ update to Governor Abbott — if the permit freeze survives it, expect more hyperscalers to route around the grid entirely.
Is the multi-agent tax a fair price for speed, or are labs shipping swarm features nobody should be paying for? Tell us in the comments.



