Anthropic opens real Claude usage data to outside researchers

Share
Anthropic opens real Claude usage data to outside researchers

The most guarded data in AI — what people actually ask a model — just became a public research resource, and the first findings are revealing. Anthropic ran a pilot that handed three outside research teams access to real, anonymized Claude usage, with almost no say over what they concluded, and one team promptly found that users are handing Claude consequential decisions.

Anthropic let researchers from Stanford, Oxford, and the safety group METR run their own studies on a quarter-million real Claude conversations, in a pilot designed to test whether independent science is possible on usage data that labs normally keep behind closed doors. The company's contractual review rights were deliberately limited to user privacy, content that could help people violate usage policies, Anthropic's confidential information, and research accuracy — everything else was the researchers' call, and they're free to publish even findings that embarrass the lab. One early result: people are delegating high-stakes tasks to Claude, a finding Anthropic says it would not have designed a study to surface on its own.

The move is bigger than a single study. Right now, real-world interaction data sits concentrated in a handful of labs, and the two options outsiders have had are lab-published analyses (which answer the lab's questions) or public datasets (which skew toward casual use). Anthropic's pilot is an attempt to open a third path: letting researchers ask their own questions of genuine production data, then publish what they find. That's the difference between the industry policing itself and the industry being studied.

Why it matters. This is a governance story, not a model story — a lab volunteering the kind of transparency regulators keep threatening to mandate. The finding that users hand Claude high-stakes tasks cuts both ways: evidence that Claude is being trusted with real responsibility, and a reminder of how little independent oversight currently exists over what frontier models are actually doing in the wild. It won't replace external audits or mandated reporting, but it's the closest thing to legitimate outside scrutiny the usage layer has seen, and the audited, anonymized setup is a template other labs could copy.

What to watch. The Stanford, Oxford, and METR papers are still being written up — their actual conclusions, once public, will tell us far more than the pilot announcement. If the data holds up to scrutiny and other labs follow suit, independent usage research could become a fixture of AI governance; if it stalls, we'll have learned how hard genuinely open data is to sustain.

Do you trust a lab to set the terms of the independent research on its own products, or should that oversight be external? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t