Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

Share
Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

Google hasn't launched Gemini 4 Argon yet, but internal documents suggest a faster follow-up is already in testing — and early impressions put it level with Anthropic's best coding model.

Google is testing a Gemini 4 variant called "Carbon" that employees say performs on par with Anthropic's Opus 5.5 on programming tasks. Business Insider, citing internal documents, screenshots, and chats, reports that Google deployed Carbon on Jetski — its internal coding platform — over the past few days, with at least one employee comparing its coding ability to Opus 5.5, Anthropic's strongest model. Carbon is one of several internal Gemini 4 variants alongside Argon (the announced frontier model) and Barium, whose Barium-B checkpoint was selected as the public Argon release. It likely ships as an update within the Argon family rather than a separate tier; one employee called it internally the "Gemini pro next model."

The timing is what makes this more than a leak. Argon only just arrived, and Carbon's arrival in internal testing within days suggests Google's model iteration loop is accelerating. DeepMind employee Vedant Misra responded to the report on X with "Have you heard of recursive self improvement" — a nod to the idea that AI is increasingly building better AI. OpenAI and Anthropic have both reported similar dynamics in their own pipelines, and Google recently mapped its lineup explicitly: Argon for frontier reasoning, Flash for speed, Omni for media, Gemma for edge. A Carbon drop would slot straight into that frontier slot before Argon even reaches general availability. We covered Argon's coding reputation before — Google employees question Gemini 4 Argon's real-world coding — and Carbon reads like the direct answer to those internal doubts.

Worth keeping the report proportionate: the Opus 5.5 comparison comes from one employee and, by Business Insider's own account, still needs more testing. Google has not commented. What to watch: whether Carbon ships as an Argon update or stands alone, and whether any of it shows up on public benchmarks before the Gemini 4 launch date lands.

Is AI building the next generation of AI models faster than labs can name them — or is one employee's chat message being over-read? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t