Coding agents as a hospital: Cockroach's five-month results

Share
Coding agents as a hospital: Cockroach's five-month results

The agentic-coding debate keeps being framed as speed versus trust. Cockroach Labs just published the most detailed trust-first ledger yet — five months of numbers, including the parts that didn't work.

Cockroach Labs ran its codebase like a teaching hospital for five months and says the pipeline merged well over a million lines of code with seven total reverts. In the setup the company calls MOLT Sinai, GitHub issues are patients, merging is discharge, and the humans in charge are Chiefs of Medicine. The core rule is that no agent writes code before its plan is reviewed: a Fellow agent reproduces the problem and posts a treatment plan, a separate Review Attending agent — running a prompt built to find fault rather than fix — approves or rejects it, and only then does work start. A Discharge Nurse at the end doesn't re-review the change at all; it verifies that the review happened properly, checking for an approval, a filled-in review template, no unresolved threads and green CI. The company says it can reconstruct why any of 1,238 pull requests merged from the issue trail alone, and that over the run it spent just over $135,000 in Claude tokens — about $84 per average issue.

The headline result is one sprint: full IBM Db2 support for its database-migration tool, built from a single GitHub issue in under two days for $4,172 in tokens. Cockroach says the equivalent Oracle work in 2024 took nine months and around $160,000 in engineering time, which its own math puts at 164 times faster and 38 times cheaper. The company is candid about the asterisk: no human reviewed the Db2 code during that initial test, and confirming it is still ongoing; since the first week, human approval has been required before anything ships to customers. The pipeline also generated much of its own work — nearly half of the 1,299 issues filed in the repo were filed by the hospital for the hospital, and 85% of the commits to the framework repo were authored by the pipeline itself. The model now runs in four Cockroach repositories.

The failure section is the part worth reading. A one-line fix went through eleven rework rounds over two days — comment hygiene, a false PR claim, rebases, flaky CI — until the safety review that documented it admitted the review loop had no circuit breaker for non-convergence; a fix that forces escalation after four consecutive rounds with new blocking findings has since merged. The instruction files rot like code: an audit of 25 skill files totalling roughly 100,000 words found 23% was removable redundancy, with the Discharge Nurse loading about 29,000 tokens of rules before reading a line of the pull request. And the bureaucracy is real — a confirmed duplicate issue still needed a human because the review protocol had no outcome for closing one.

The claim that matters isn't the speed, it's the metric swap. Most agent-automation experiments optimize pull requests per hour; Cockroach explicitly asked how many of its merges the team "would have been embarrassed to have merged ourselves," which is why the pipeline is slow, layered and expensive on purpose. An industry roundup this month had Cockroach's co-founder arguing that code review is heading toward obsolescence — the hospital paper is the evidence file for that argument, warts included. The open wound the authors name themselves: the hospital doesn't teach. Junior engineers don't get the judgment reps that writing and defending plans build, and their proposed fix — humans writing plans that the agents critique — isn't built yet.

What to watch: whether the framework escapes Cockroach's four repos, and whether any competitor publishes a matching ledger with a human-reviewed baseline.

Would you ship agent-written database code on a seven-revert track record — and who signs off when the reviewer is also an agent? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t