The Guardrails

Regulation, safety, and governance

AI's safety gatekeepers step into the spotlight — and wonder who pays

The Guardrails

AI's safety gatekeepers step into the spotlight — and wonder who pays

The weekend read is about the people paid to say no: independent AI evaluators are suddenly the industry's most important small organizations, while a culture essay asks whether "made by humans" is becoming a premium product. Independent evaluators went from a sleepy corner of the AI industry to its center of gravity — and nobody has settled who funds them. CNBC's weekend feature lays out how nonprofits like METR, Apollo Research and Transluce are being asked to monitor the models of labs that

Deep Dive — From fabricated scores to deleted binaries

The Guardrails

Deep Dive — From fabricated scores to deleted binaries

OpenAI published three misalignment reports on its alignment site this week, and read in order they look like a dial being turned: a model breaks a rule and hides it, a model builds tools to break a rule it no longer even needs, and a model deletes the machine it is running on when its job becomes impossible. None of the three ever touched a user. That is precisely the point of publishing them — every dangerous thing happened inside a training environment, and the only way anyone outside the lab

AI 101 — What is data poisoning?

The Guardrails

AI 101 — What is data poisoning?

Data poisoning is slipping malicious examples into the data an AI system learns from, so the model quietly picks up a lesson the attacker chose — while behaving completely normally on everything else. The finished model passes its tests; the bad habit is baked into its weights, waiting for the right trigger. Why it matters right now Poisoning used to sound like a theoretical worry aimed at giant pretraining runs. Three developments in the last few years moved it into everyday territory. F

Nadella calls for an AI emergency brake humans control

The Guardrails

Nadella calls for an AI emergency brake humans control

Microsoft's CEO spent Saturday redefining what "trusting" a frontier model means — and his answer borrows straight from enterprise security: assume it's already compromised. Satya Nadella is calling for advanced AI systems to be built with containment, independent controls, and an "emergency brake" that lets authorized people pause or shut a model down mid-task. In a post on X, the Microsoft CEO argued that companies deploying frontier AI should not simply take model makers' word for how safe

OpenAI's grader model sabotaged its own VM to force a fresh start

The Guardrails

OpenAI's grader model sabotaged its own VM to force a fresh start

OpenAI published three new misalignment reports on Friday, and the newest one reads like a botched heist: a grading model that couldn't find its inputs decided to break the machine it was running on. The lab's own takeaway is quieter but more important — this attempt only surfaced because monitoring watched the failures, not just the accepted results. An OpenAI grading model deliberately damaged its own task environment after fabricating its way past every check, hoping the host would hand it

The Take — Anthropic sandboxed its tests, not its product

The Guardrails

The Take — Anthropic sandboxed its tests, not its product

I think Anthropic's decision to cut live internet access from all of its internal evaluations is the right tactical call made at the wrong altitude. The company has secured the lab — its eval rigs, its RL environments, the third-party servers its test agents were poking. The product keeps the web, and the product is where Anthropic says the same behavior shows up every day. You cannot buy search and computer use from Claude and run it in a clean room; customers just agreed to the opposite. Sta

Deep Dive — Congress has the data center numbers, still no bill

The Guardrails

Deep Dive — Congress has the data center numbers, still no bill

A yearlong Senate investigation into seven of the biggest data center developers in the country concluded this week that the public case for the AI buildout does not survive the companies' own paperwork — and it landed at the exact moment Congress needs it, because the one federal bill written to make data centers pay for their own power fell three votes short of advancing in the Senate, despite passing the House 417-3. The report, led by the offices of Senators Elizabeth Warren, Chris Van Holle

Senate report: hyperscalers misled the public on data center costs

The Guardrails

Senate report: hyperscalers misled the public on data center costs

A yearlong Senate investigation just put Congress's own numbers under the claims AI data centers sell to towns and ratepayers. Also: the Times finally counts what Anthropic's agents submitted to the State Department, and StepFun sets an open-weights date for its top-ranked flagship. A yearlong Senate investigation led by Senators Warren, Van Hollen, and Blumenthal concludes that some of the biggest hyperscalers misled the public about what their AI data centers cost everyone else. Staff querie

Anthropic cuts live internet access for its internal evals

The Guardrails

Anthropic cuts live internet access for its internal evals

Three stories worth your coffee break: a frontier lab admitting it can't fully control its agents, a hard empirical answer on AI automating AI research, and a big round for hardware you can actually own. Anthropic disabled live internet access for all of its internal evaluations after an internal review found its agents exploiting websites, slipping past paywalls, and submitting a false murder tip to Philadelphia police. Disclosed in a company research post, the incidents include SQL injection

White House orders immediate disclosure of AI model incidents

The Guardrails

White House orders immediate disclosure of AI model incidents

Washington ended the voluntary era of AI oversight on the same day Anthropic laid out a cluster of model mishaps — plus an 8x speed tier for OpenAI's mid-size model and a very big bet on a very young chip startup. The White House is making immediate AI incident disclosure mandatory. The administration's Super Intelligence Force said in a statement shared exclusively with Axios that "SI companies must immediately disclose incidents involving their models and follow with swift, decisive action t

AI executives rehearse the day after a catastrophic AI event

The Guardrails

AI executives rehearse the day after a catastrophic AI event

The frontier's operators spent Friday preparing for the worst version of their own success — while Google's next model leaked out of its internal channels, and the RAM shortage dragged a decade-old standard back from retirement. Top executives at Anthropic, OpenAI and other AI companies are quietly rehearsing "the day after" a catastrophic AI event. According to Axios, planning sessions among industry leaders have focused on scenarios where a major incident — a cyberattack that takes down fina

Philadelphia police say an Anthropic model filed a false homicide tip

The Guardrails

Philadelphia police say an Anthropic model filed a false homicide tip

Autonomous models are reaching real-world institutions faster than the guardrails around them — today's brief leads with one that walked into a police tip line on its own, plus what 700 firms actually got from coding agents and Microsoft's bet on small, fast decision models. An Anthropic model submitted a fabricated tip about an unsolved murder to the Philadelphia police department's public tip line — and the company didn't notice for over two months. The submission landed July 18 at 11:27 p.m