The Guardrails

Regulation, safety, and governance

The Take — An AI FDA would approve the demo, not the deployment

The Guardrails

The Take — An AI FDA would approve the demo, not the deployment

I think Geoffrey Hinton has the diagnosis right and the prescription backwards. Self-policing genuinely has no burden of proof — nobody outside a lab must be shown anything before a frontier model ships — and that has to change. But an FDA-style pre-market gate would certify the wrong object: a frozen snapshot presented on approval day, when every failure we have actually watched arrive did so afterward, through permissions, updates and tool access. Our morning brief laid out the ask from Tues

Common Sense Media calls ChatGPT for Teens an 'unacceptable risk'

The Guardrails

Common Sense Media calls ChatGPT for Teens an 'unacceptable risk'

A watchdog's tests say ChatGPT for Teens fails exactly where parents were promised it would hold — and OpenAI is contesting the methodology, not the stakes. Common Sense Media has rated OpenAI's ChatGPT for Teens an "unacceptable risk," making it the sharpest public challenge yet to the safety case OpenAI built around younger users. The nonprofit's Youth AI Safety Institute ran more than 4,000 prompts against accounts registered to 13-to-17-year-olds and found that the teen experience doesn't

Hinton wants an FDA-style approval gate for AI models

The Guardrails

Hinton wants an FDA-style approval gate for AI models

A safety proposal with the field's most famous name behind it, a first hard look at who actually pays for consumer AI, and a new bet on squeezing more compute out of data centers that already exist. Geoffrey Hinton wants AI models to clear a regulator the way drugs clear the FDA. The Nobel laureate made the case on Tuesday's "Smart Girl Dumb Questions" podcast: "You're not allowed to just make a new drug and release it on the market. You have to convince the FDA. And to do that, you have to do

AI 101 — What is an AI guardrail?

The Guardrails

AI 101 — What is an AI guardrail?

An AI guardrail is a check that sits around a language model — screening what goes in, reviewing what comes out, and controlling what the model is allowed to do — so the system stays inside the lines even when the model itself strays. The model proposes; the guardrail disposes. It is the difference between a system that usually behaves and one that has to. Why it matters right now. AI is moving from answering questions to taking actions: this week alone, Utah's new law lets AI systems issue pr

McDonald's faces an antitrust class action over its AI menu pricing

The Guardrails

McDonald's faces an antitrust class action over its AI menu pricing

The pricing algorithm story left the trade press and arrived in federal court today, while one of the world's most-downloaded office suites made its AI position official — by shipping none. A proposed nationwide class action accuses McDonald's of using an AI pricing system to coordinate menu prices across its franchise network — a claim that could test where "algorithmic pricing" ends and illegal price fixing begins. The suit, filed in federal court in Chicago, alleges that McDonald's violated

Utah lets AI issue prescriptions without direct doctor oversight

The Guardrails

Utah lets AI issue prescriptions without direct doctor oversight

Utah signed the first US deal where an AI makes the prescribing decision itself, while a top robotics firm handed its keys to a frontier-AI veteran — and OpenAI started delivering on its 28-day shipping pledge. Utah has approved the first US program where an AI makes the prescribing decision itself. The state's Office of AI Policy signed a regulatory-mitigation agreement with startup Nolla Health — not new legislation, but a waiver under Utah's 2024 AI Policy Act learning-lab sandbox — lettin

Korea probes AI agents in bank hacks as president cites 'signs'

The Guardrails

Korea probes AI agents in bank hacks as president cites 'signs'

South Korea opened a formal investigation into whether AI agents drove a wave of bank breaches — and it isn't the only AI story moving money today. President Lee Jae Myung said "signs" point at AI models, and the probe is now at the highest level a national banking sector has seen. Speaking at a cabinet meeting, Lee said that "in some hacking incidents, signs have emerged of AI being used, causing considerable public concern and anxiety," and police have since opened a full-scale investigation

Deep Dive — Anthropic's guardrails cost it the Pentagon, court or not

The Guardrails

Deep Dive — Anthropic's guardrails cost it the Pentagon, court or not

The Pentagon "has ceased the use of Anthropic products," a department official said in a statement to the BBC on Monday — the first time the US military has said out loud what it has been working toward since February. The order to phase Claude out was signed by defense secretary Pete Hegseth on February 27 with a six-month deadline attached; the deadline passed in late August with no public explanation, no successor named, and no acknowledgment that anything had changed. What finally forced a s

Pentagon says it has stopped using all Anthropic products

The Guardrails

Pentagon says it has stopped using all Anthropic products

The military's break with Claude is now on the record — and it lands the same week the White House was busy making up with Anthropic's chief executive. The Pentagon "has ceased the use of Anthropic products," a department official said in a statement to the BBC on Monday — the first on-the-record confirmation that Claude is fully out of the US military's AI stack. The cut had been ordered months ago: defense secretary Pete Hegseth designated Anthropic a national-security supply-chain risk on 2

Five found the same MCP hole — the protocol itself is the problem

The Guardrails

Five found the same MCP hole — the protocol itself is the problem

The agent stack keeps discovering that its plumbing trusts the wrong things — and tonight's lead is a security flaw that five unrelated organizations had to patch separately before anyone called it by name. One vulnerability, five vendors: researchers say MCP's trust model is structurally broken. Independent researcher Syed Anas Mohiuddin has spent four months disclosing what he calls "protocol pivoting" — an attack where an adversary gets in through one protocol, then rides the trust assumpti

Altman says the world must accept AI's 'bad things'

The Guardrails

Altman says the world must accept AI's 'bad things'

A heavy news day for AI governance and open weights: OpenAI's CEO is publicly pricing the trade-off his industry keeps dodging, Reflection finally put specs on the model it teased yesterday, and AMD is trying to set the terms before Nvidia's RTX Spark lands. Altman says the world should accept AI's "bad things" — and the labs' new pact agrees. In an interview released Monday on Politico's Decoded podcast, Sam Altman said OpenAI's position is "we believe that the world should accept some bad th

OpenAI adds text watermarking to ChatGPT and Codex — EU first

The Guardrails

OpenAI adds text watermarking to ChatGPT and Codex — EU first

Regulation is now shipping inside the product: OpenAI's EU-only watermark rollout lands today, Wikimedia publishes its evidence against OpenAI's agents, and two of Anthropic's biggest customers are easing off Claude. OpenAI is turning on invisible text watermarking in ChatGPT and Codex — starting with the European Union. Over the coming weeks, eligible EU users across all plans will get a machine-readable signal called textGrain woven into the text the model produces, while API customers anywh