Anthropic launches cyber defense program and free OSS Scanner

Share
Anthropic launches cyber defense program and free OSS Scanner

The labs keep sliding into security work: Anthropic put eleven marquee security partners behind critical-infrastructure defense, a benchmark found the best model still can't rebuild two-thirds of ordinary programs, and AMD is promising much more silicon for 2027.

Anthropic has launched the Anthropic Cyber Mission: a Critical Infrastructure Defense Program pairing its models and engineers with eleven founding security partners, plus a free vulnerability-scanning service for open-source projects called OSS Scanner. The partner roster is the story — Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation — and the target is operational technology: the substations, water systems and plant controls that run for decades and often cannot be taken offline for a patch. Booz Allen's Andrew Turner called operational technology "the next frontier for autonomous AI-enabled attacks," with the open question being how much control AI can gain over an industrial process and how quickly. The commercial terms are conspicuously absent: Axios noted Anthropic hasn't said whether partners get free model access or who covers the compute costs, and how consultants will test fixes without disrupting live utility operations remains unanswered.

The OSS Scanner half is the more novel experiment. Anthropic says its models surfaced more than 29,000 candidate vulnerabilities in widely used software over six months; staff manually reviewed about 6,000, and roughly 5,000 reports — several with proposed patches — have gone out to maintainers. OSS Scanner turns that into a standing service, and reports go straight to maintainers without human review so they arrive faster, mistakes and all; Anthropic expects a true-positive rate above 90% and backs it with an audit in which expert penetration testers cleared 85 of 97 critical and high-severity findings for disclosure, while wolfSSL reported all but two of its 74 early-trial reports were valid and five became CVEs. Our read: the interesting number is 29,000 found versus 6,000 reviewed — discovery is automating faster than verification can keep up, so the bottleneck (and the risk) moves to maintainers deciding which machine-generated bug reports to trust.


A new benchmark called Behavior2Code says frontier models are still bad at the most human of programming skills: understanding software they can only see the behavior of. Turing's Frontier Research Lab hides the source and asks six models to rebuild 60 real command-line programs from how they behave — 1,080 runs, 6,872 graded assertions across 10 languages. The best result was GPT 6 Astra at 19 of 60 programs after three attempts; Claude Opus 5 and Opus 5.5 both managed 14, Grok 4.6 got 7, Gemini 3.7 Flash got 1, and GPT-5.6 Sol scored zero. Thirty-five programs defeated every model tested — and the misses are nearly-there: Claude Opus 5's best attempt at the cmark library passed 749 of 756 checks before failing on the last few.


AMD chief Lisa Su says her company will "substantially increase our supply in 2027" — and is now planning wafer capacity three to five years out instead of one to two. Su told reporters in Taipei this week that demand will stay "very, very high for the next several years" as she met Foxconn, was due to sit down with TSMC, and headed to Korea to press memory makers — with AMD's market value having recently topped $1 trillion. No volume figures came with the pledge, but the direction matches what Nvidia's Jensen Huang has been saying: AI inference demand is outrunning what the supply chain can currently ship, and the fix is multi-year commitments made now.

What to watch: whether Anthropic publishes the commercial terms — and partner retention — for its infrastructure program, and whether any model cracks Behavior2Code's remaining 35 unsolved programs.

Should AI labs be allowed to send machine-generated vulnerability reports to maintainers without human review, or does that just dump the verification burden on volunteers? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t