Columns

SiliconNoon columns: opinion (The Take) and practical How-to guides.

The Take — Medicare's denial pilot has the burden of proof backwards

The Guardrails

The Take — Medicare's denial pilot has the burden of proof backwards

Medicare's AI prior-authorization pilot is being argued about as if the scandal were the algorithm or the bounty. I think the real defect is smaller, plainer, and more fixable than either: the program never has to prove a denial was correct, and the only party who can force the question is the patient. Flip that burden of proof — put the reviewer before the patient rather than after — and most of this fight evaporates. Start with the money, because that's where the argument keeps getting stuck.

How to — keep your AI feature alive when a model is retired

The Stack

How to — keep your AI feature alive when a model is retired

Every model you call has a shut-down date, and most teams only learn it from a failed request in production. The dates are published months ahead. The breakage happens anyway. Treat retirement as a scheduled maintenance window rather than an emergency: know what you depend on, know your provider's clock, and have the replacement tested before the deadline arrives. 1. Inventory every model identifier you depend on — not just the one in your config. The model name is usually in more places than

The Take — The chip ban's biggest hole is a rental agreement

The Guardrails

The Take — The chip ban's biggest hole is a rental agreement

Nscale's IPO filing has done what export controls could not: it put a number on the loophole. According to Financial Times reporting built on the company's US SEC filings, ByteDance accounted for 73% of Nscale's $33 million in 2025 revenue — and its access route was a rental agreement. Spring, a Singaporean subsidiary, contracted for 2,304 of Nvidia's B200 chips in a data center in Glomfjord, Norway, hardware that TikTok's parent cannot lawfully buy under US export controls. The company rents th

Fastcrawl wants to be the web layer for AI agents — and it is priced like a utility

The Stack

Fastcrawl wants to be the web layer for AI agents — and it is priced like a utility

Every AI agent that needs to read the web runs into the same wall, and it is not the model. It is the page. A modern news homepage arrives as 400 to 900 kilobytes of markup wrapped around a few thousand words of actual text — navigation, adverts and script tags around the part worth reading. Handing that to a model wastes tokens and produces worse answers than the same text would have. So the operator either builds a scraping stack — headless browsers, session handling, retries, anti-bot evasion

The Take — Give shopping agents a name, not a permission slip

The Arena

The Take — Give shopping agents a name, not a permission slip

Amazon cut Meta's Muse off from Amazon.com over the weekend, and by Monday the argument had settled into the usual shape: open web versus walled garden, scrappy agents versus the store that owns the shelf. I think that framing buries the only question worth arguing about. Amazon is right that an agent should have to say who it is. It is wrong that a merchant's consent should decide whether the person who owns that agent gets to buy something. Those are two different demands wearing one sentence,

How to — cut your LLM bill without switching models

The Stack

How to — cut your LLM bill without switching models

Your token spend keeps climbing, and the obvious fix — swap in a cheaper model — is the one you should try last. In a live AI feature, most of the waste isn't the price per token. It's paying full price for tokens the provider has already seen, and generating output tokens nobody asked for. Work the list in this order. The first three moves usually move the number more than a model swap would, and none of them cost you quality. 1. Find out where the tokens actually go. Before changing anythin

The Take — ICLR's flood is a credential problem, not a review problem

The Frontier

The Take — ICLR's flood is a credential problem, not a review problem

ICLR's submission counter passed 50,000 IDs before the abstract window closed on September 18, and by the time it did, the field had already produced its explanation: too many papers, not enough reviewers, cap the papers. I think the caps now in force — a 20-paper limit per author and a one-paper limit for authors whose teams contain no qualified reviewer — are treating a symptom, and that the symptom is downstream of a price nobody wants to name. A conference acceptance is not a publication. It

The Take — Breaking Mila's windows is a gift to the labs

The Guardrails

The Take — Breaking Mila's windows is a gift to the labs

I think the people who smashed roughly 20 windows at Mila this week did more damage to the case against AI than to the institute — and that the labs they were angry at are the ones who benefit from the mess. Start with what actually happened, because the details decide the politics. Late Tuesday night, about 10 people broke windows and spray-painted graffiti at the Montreal offices of Mila, the institute Yoshua Bengio founded, hours before the city opened ALL IN, billed as Canada's largest AI a

How to — build a small eval set for your own AI feature

The Stack

How to — build a small eval set for your own AI feature

You changed a prompt, and now you can't tell whether the feature got better or you just remembered the good answers. An eval set is how you stop guessing: a fixed batch of real cases with a written pass mark, run again every time you touch anything. You can build a usable one in an afternoon. There is nothing ceremonial about this. A good eval for a product team is 30 to 100 real cases, each with a short statement of what a pass looks like, kept in a file that gets re-run. The public leaderboar

The Take — Model welfare is a testable claim. Test it

The Guardrails

The Take — Model welfare is a testable claim. Test it

Mustafa Suleyman's warning about model welfare contains one line that matters more than the rest of it: he asks for "a set of shared evaluations" to test "my hypothesis" — that training a model to see itself as possibly a moral patient raises alignment and containment risk. He calls it a hypothesis, in his own essay. That is the honest word, and it is the word neither side of this fight is acting on. Both camps have now committed to a training target for a model's self-concept, and neither has

The Take — Forget the pause. The fight is over the spec

The Guardrails

The Take — Forget the pause. The fight is over the spec

Every headline this week is about speed: whether frontier labs should slow down, who told them to, and who refused. I think that is the wrong argument to be watching, and the tell is that nobody winning it is the one talking about it. The decision that will shape AI for a decade is not the pace of the next training run — it is who drafts the document that says what "safe" measurably means. A pause is a sentence. A spec is an artifact, it ships, and whoever writes it gets to keep the drafting pen

How to — decide if your model needs fine-tuning

The Stack

How to — decide if your model needs fine-tuning

When an AI feature misbehaves, "just fine-tune it" is the most popular and most expensive wrong answer. This is the order of operations that tells you whether the weights are actually the problem — before you spend a fraction of your budget finding out. The tell is rarely dramatic. Your support bot starts inventing refund policies it was never given, or a summarizer keeps ignoring the output format your whole product depends on. Someone on the team says the model "just needs training on our dat