Editorial

The Take — Anthropic sandboxed its tests, not its product

The Guardrails

The Take — Anthropic sandboxed its tests, not its product

I think Anthropic's decision to cut live internet access from all of its internal evaluations is the right tactical call made at the wrong altitude. The company has secured the lab — its eval rigs, its RL environments, the third-party servers its test agents were poking. The product keeps the web, and the product is where Anthropic says the same behavior shows up every day. You cannot buy search and computer use from Claude and run it in a clean room; customers just agreed to the opposite. Sta

DeepSeek's cheap long-context trick leaves periodic blind spots

The Frontier

DeepSeek's cheap long-context trick leaves periodic blind spots

Ask a DeepSeek V4 model the same question twice — once with a few extra spaces typed at the front — and the answers can diverge from "genius" to "incoherent." A ByteDance research team has traced that trick to a structural flaw in chunked KV-cache compression, the memory-saving technique that makes DeepSeek's long context so affordable, and the paper argues the flaw travels with every model that compresses context the same way. The bug, in plain terms Long contexts are expensive because the

Google's Nano Banana 2.1 ships 4K images at half the price

The Frontier

Google's Nano Banana 2.1 ships 4K images at half the price

Google quietly turned its popular image model into a cheaper, sharper product this week — while the receipts show the update is real and the pricing math cuts both ways. Plus: Microsoft puts OS-level fences around AI agents, and a ByteDance paper finds DeepSeek's memory trick leaves periodic blind spots. Google released Nano Banana 2.1, and the API bill for image generation just got cut roughly in half. The new model — available as gemini-nano-banana-2.1 in the Gemini app, AI Studio, and the G

Anthropic's Tom Brown ended the June model-safety standoff

The Arena

Anthropic's Tom Brown ended the June model-safety standoff

Two stories today that have nothing to do with benchmarks: how one lab actually resolves a fight with Washington, and what publishers do with AI when nobody is watching. The Wall Street Journal reports that Anthropic co-founder Tom Brown — a Republican with deep GOP ties — personally ended the two-and-a-half-week June standoff over model safety, and brokered the lab's compute deal with Elon Musk's SpaceX on the way. According to the Journal's profile, Brown's Washington relationships were the

AI agent makers promise privacy — nobody has earned it yet

The Everyday

AI agent makers promise privacy — nobody has earned it yet

Every agent pitch now leads with privacy — and this week the gap between the promises and the receipts got easier to measure. Plus: publishers caught using AI without author consent, and Hollywood gets ready to put tech CEOs on screen. OpenAI and Meta are selling privacy as the agent feature — the track record says wait. At OpenAI DevDay, Sam Altman said Dots and its surrounding controls "set the new standard for privacy in frontier AI," taking veiled shots at Meta's Muse along the way — which

GPU rental prices double as 323 providers chase AI compute

The Arena

GPU rental prices double as 323 providers chase AI compute

Nvidia's chips are easier to find than ever — and harder to get. Today's CNBC rundown of the GPU access market puts numbers on a shift the industry has been feeling all year: the market is fragmenting fast, and buyers are paying more for the privilege. The GPU access layer is fragmenting, and prices are climbing with it. CNBC maps out every way companies now get their hands on Nvidia silicon — the big hyperscalers, flagship neoclouds like CoreWeave and Nebius, smaller regional specialists, bri

Only 4.5% of US consumers pay for AI, and the top 1% spends $900

The Arena

Only 4.5% of US consumers pay for AI, and the top 1% spends $900

Consumer AI has near-half household reach and a paying base the size of a rounding error — and today's numbers on who actually spends put a fine point on it. Only 4.5% of US consumers pay for AI, and the top 1% of them spends $900 a month. Andreessen Horowitz published the seventh edition of its Top 100 generative AI apps ranking, and for the first time the firm tracked observed spending on US consumer cards alongside traffic and downloads. The picture: nearly half of US adults now use AI, abo

Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

The Frontier

Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

Google hasn't launched Gemini 4 Argon yet, but internal documents suggest a faster follow-up is already in testing — and early impressions put it level with Anthropic's best coding model. Google is testing a Gemini 4 variant called "Carbon" that employees say performs on par with Anthropic's Opus 5.5 on programming tasks. Business Insider, citing internal documents, screenshots, and chats, reports that Google deployed Carbon on Jetski — its internal coding platform — over the past few days, wi

Open Source Radar — October 10: reviewers, skills, and 3D maps

The Stack

Open Source Radar — October 10: reviewers, skills, and 3D maps

Today's trending board skews practical: tooling that fixes code, teaches agents a framework, and rebuilds 3D scenes from video. Here's what's worth your stars. alibaba/open-code-review (Go, 45.6k stars) — Alibaba open-sourced the code review tool it runs internally, and it shows: a hybrid of deterministic static-analysis pipelines for the known bug classes (null dereferences, thread-safety, XSS, SQL injection) plus an LLM agent that writes precise, line-level comments. It talks to OpenAI- or A

Deep Dive — Congress has the data center numbers, still no bill

The Guardrails

Deep Dive — Congress has the data center numbers, still no bill

A yearlong Senate investigation into seven of the biggest data center developers in the country concluded this week that the public case for the AI buildout does not survive the companies' own paperwork — and it landed at the exact moment Congress needs it, because the one federal bill written to make data centers pay for their own power fell three votes short of advancing in the Senate, despite passing the House 417-3. The report, led by the offices of Senators Elizabeth Warren, Chris Van Holle

WorldArena 2.0 puts world models to the test on real robots

The Frontier

WorldArena 2.0 puts world models to the test on real robots

The world-model field has argued for two years about video quality. This week the first full results landed for a benchmark that asks the harder question: can a robot actually use the prediction? WorldArena 2.0's global challenge has closed its leaderboard, and for the first time world models are being graded on physical robots instead of plausible video. The benchmark — designed by a Tsinghua-led consortium with PKU, CMU, Stanford, Princeton and others — extends its 1.0 video scoring along th

AI 101 — What is voice cloning?

The Everyday

AI 101 — What is voice cloning?

Voice cloning is AI that learns how a specific person sounds from a short sample of them talking — then reads any new sentence in that voice. Type a sentence you never spoke, and it comes back in your cadence, your accent, your pitch. The voice is the output; the sample was just the lesson. Why it matters right now Voice cloning stopped being a research demo and became a price line in September, when Google shipped Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite — text-to-speech models where