OpenAI's dots go mobile and can start your Codex work

One product update and two benchmark papers this hour, all circling the same question: how much of the work can you actually hand over, and what happens to the part you didn't?
OpenAI's dots — the always-on agents that get their own cloud computer — can now be created and configured entirely from the ChatGPT app on iOS and Android, and a dot can start Codex work or pick up an existing Codex thread on your behalf. The October 9 release, billed as the first of a fresh run of updates for dots, moves setup off the desktop: you name the dot, set its appearance, connect plugins and make it the opening conversation from your phone, and the dot decides whether to continue an existing Codex thread or start fresh, drawing on your ChatGPT conversations, Codex history and automations. Dots can also read and edit ChatGPT Work automations, which makes them less a chat window with a computer attached and more an operations layer sitting over the rest of the product. Access is still gated — Pro users outside the EEA, Switzerland and the UK, Business Premium, and Enterprise workspaces where an admin switches it on — so the real question is whether the phone gives dots enough surface to become OpenAI's front door to agents rather than a companion feature. We covered the original launch — OpenAI's Dots give Pro users an agent with its own computer.
When AI agents work as a team, the best model in the group finishes only 52% of the tasks — and less than a third of what the team does actually moves the task forward. AgentWorld, a benchmark accepted at COLM 2026 from OpenAgents, Columbia, Penn, Seoul National and Penn State, runs 100 hand-written tasks plus 100 variants inside an MMORPG-style sandbox, with teams of 3 to 20 agents that cannot see each other's internal state and must plan, talk and share resources to win. The best of the four models evaluated — Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B — cleared 52% of tasks, and the paper's causal-collaboration score of 0.320 says under a third of team actions contributed to the outcome. Read the roster carefully: these are Flash/Mini/Haiku-tier models, not flagships, so this is the price tier most multi-agent deployments will actually run at — and at that tier half the work lands.
AI agents taking on quantum-engineering lab tasks top out at 78%, and their most common failure isn't a wrong answer — it's never submitting one. A preprint from researchers at the National University of Singapore, NTU and the Unitary Foundation introduces Quantum-Harbor and its QIQCBench eval: 49 tasks run against 17 agent systems from 8 vendors, with 117 trials per system. Pass rates span from GPT-5.6-sol's 78% down to GLM-5.2 at 3%, with Claude Opus 4.8 at 51%; of 1,510 recorded failures, 937 were the agent never submitting anything at all, and seven systems timed out before submitting in more than 60% of trials. "Agents can demonstrate a task and still fail to run it autonomously. That is the gap we found," co-author Shihao Ru said of the work, which is not yet peer reviewed — but the failure mode it measures is the one that bites in production.
What to watch: whether OpenAI widens dots beyond its gated tiers during its run of daily updates, and whether AgentWorld's collaboration score becomes a standard benchmark line next to raw task success.
If fewer than a third of your agents' actions actually contribute, is a team of them a workforce or an audience? Tell us in the comments.



