HiDream-O1-Embodied tops RoboColiseum's robustness chart

Share
HiDream-O1-Embodied tops RoboColiseum's robustness chart

A Chinese world-model lab's first move into embodied robotics lands on top of the toughest sub-score, while Alibaba ships a "multi-user workbench" that turns a sentence into a working group app.


HiDream-O1-Embodied debuted at number one on the RoboColiseum robustness leaderboard, scoring 0.692 in the benchmark's most adversarial sub-track on its first submission. HiDream.ai (智象未来) framed the release as the "real-world" leg of a world-model strategy that already covers image (HiDream-O1-Image) and interactive video (HiDream-O1-World, which topped the WBench Navi sub-leaderboard a month ago). The new model closes the loop: the interactive world-model handles "understand and predict," the embodied one handles "operate and execute." RoboColiseum's robustness track is the hard one — it changes background, lighting, materials, camera position, image quality and rewrites the instructions, then checks whether the policy still works in a "greenhouse" the lab never saw. CTO Yao Ting tied the score to three engineering bets: instruction understanding that survives paraphrasing, multi-camera visual fusion that holds up when one view is occluded, and a training pipeline that feeds the model noisy, partially corrupted inputs on purpose. The data side is the moat. HiDream uses mocap-grade human motion from partner Noitom as the "real base," then has the model itself generate the long tail of background, lighting and object variations — making the model both "examinee" and "examiner." For a Chinese lab that has been quietly racking up image and world-model wins since last summer, an embodied number-one on a tough benchmark is the strongest signal yet that "native omnimodal" — one stack from pixels to actions — is a real alternative to bolted-on visuomotor policies.


Alibaba's Qianwen Office shipped what it calls the industry's first "multi-user workbench": type a sentence, get a hosted web app that up to a hundred people can run at once with role permissions, a backend and cloud storage. Most AI workspace tools stop at single-user generators — a resume page, a personal check-in tracker, a portfolio. Qianwen's version adds four enterprise-grade pieces: roles (admin vs. member vs. guest), a shared database, a management console, and one-click publish to a shareable link. Use cases the team demoed read like a roll-call of friction every small org already knows: a market organizer running vendor sign-ups, deposits and booth assignments for 100 stalls; a teacher building a homework and grading tool that students and parents each see only their own slice; a brand managing influencer and supplier workflows across cities. The pitch — explicit in Alibaba's own copy — is that you can replace a ¥10,000-a-year vertical SaaS contract (or a six-figure custom dev cycle) with a natural-language prompt, then keep editing the workflow in plain language as the business changes. Qianwen Office says it crossed 30 million users in its first month, with enterprise accounts over half of that base. The interesting bet is the positioning: Alibaba is not selling "AI for the individual knowledge worker," it is selling "AI that builds your internal tools." If the workbench holds up, every other Chinese agent suite now has a "build a working app from a sentence" reference point to match or beat.


What to watch: HiDream has topped image and interactive-world leaderboards before, but a single embodied win is one data point — the real question is whether robustness holds when real robots run the policy outside simulation. For Qianwen Office, the test is whether 30 million sign-ups translate into paying teams, or whether the workbench is a productivity demo that does not survive the second business process.

Which of these two — embodied world models or agent-built internal tools — do you think will reshape Chinese enterprise AI first? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t