Microsoft builds its decision model on Qwen, not OpenAI

Share
Microsoft builds its decision model on Qwen, not OpenAI

Sunday's haul mixed one real model launch with a milestone for China's driverless delivery fleets — and a robotics round that's really about drug-making capacity.

Microsoft launched Microsoft-Decision-1, a fast decision-scoring model built by post-training Alibaba's Qwen3.5-9B. Unlike an LLM that generates text, a decision model returns a calibrated probability — the kind of call an agent makes when it routes a ticket, flags an incident, or approves a refund. Microsoft says the model takes first place in its own 36-benchmark comparison covering nearly 150,000 questions, answers at 85 milliseconds at the median, and costs $0.042 per million input tokens with output free. It's available now in Foundry and on OpenRouter, and Microsoft frames it as the first in a new class of models it will soon rebase on its own MAI models and on OpenAI's.

Every headline number is Microsoft's own, and none have been independently verified — including the claim that it runs 35 times faster at median latency than GPT-6 Sol. H2O.ai has already disputed Microsoft's latency comparison, arguing that adjusted figures for its own model double from 29 milliseconds to 210 milliseconds, while Microsoft's model doesn't appear on the JevBench leaderboard at all. The interesting part isn't the leaderboard, it's the base model: Microsoft shipped a flagship release on an open-weight model from a Chinese lab, which tells you Qwen still powers products inside its rivals.


Neolix now runs what Bloomberg calls the world's largest driverless delivery fleet — 27,000 robovans across 300 cities in 15 countries. The Beijing company says that's nearly five times the combined size of Waymo's and Baidu's robotaxi operations, and it's targeting 50,000 domestic vehicles this year. The operational milestone is Shenzhen, which opened nighttime driverless delivery in March 2026: after-dark routes there have grown from 2 to 437, with more than 170 vans moving parcels overnight while 42 vehicle metrics stream into a municipal monitoring platform. One caution — the 27,000 figure is Neolix's own, and no outlet outside Bloomberg has independently verified it.


Multiply Labs raised $75 million to robotize drug manufacturing — a Series B led by Patrick Soon-Shiong's NantWorks, with AstraZeneca, Lingotto and Teradyne joining. The company's enclosed robotics clusters imitate human manufacturing steps, so the output doesn't require fresh regulatory approval; Multiply claims a 100x throughput boost and 74% lower cost per dose. CEO Fred Parietti's framing is the story: "AI is designing more therapies than the industry will ever be able to manufacture." With more than $100 million raised to date, the bet is that the AI era's bottleneck is physical capacity, not model capability.

What to watch: whether the promised MAI- and OpenAI-based versions of Decision-1 actually ship, and whether its latency claims survive independent benchmarking now that H2O.ai is contesting them.

Is a 9B model that scores decisions cheaper and faster than a frontier LLM the more meaningful release? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t