Claude Haiku 5.5 arrives at up to 90% lower prices

OpenAI spent today putting GPT-6 and its interactive answers into ChatGPT — we covered that this afternoon in OpenAI launches Intelligent UI, making ChatGPT answers interactive. Anthropic's answer, hours later, was not a bigger model but a cheaper one.
Anthropic released Claude Haiku 5.5, its cheapest and fastest small model yet, and cut prices across the board. For prompts up to 100,000 tokens — about 90 percent of all traffic to the previous Haiku — the new model costs up to 90 percent less than Haiku 4.5, and half as much above that threshold. The benchmark jumps are large where cheap models matter most: computer use on OSWorld 2.1 goes from 15.7 percent to 72.4 percent, agentic coding on Terminal-Bench 4.0 from a flat zero to 39.2 percent, and Humanity's Last Exam from 10.2 to 45.9 percent without tools. Anthropic's own comparison table has Haiku 5.5 beating OpenAI's budget GPT-6 Luna in every tested category, and it's the first Haiku-class model with adjustable effort settings, so callers can trade intelligence for cost per request.
The launch is really a pricing move. Alongside Haiku 5.5, Anthropic halved cache-read prices on Sonnet 5.5 — $0.10 per million tokens instead of $0.20 — which it says makes agentic work about 20 percent cheaper, and is adding monthly API credits for subscribers ($100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team plans). That lands a day after OpenAI's GPT-6.1 pricing, and The Decoder reads it as the price war hardening rather than easing. One caveat worth holding onto: Haiku 5.5 ships with an updated tokenizer that uses slightly more tokens per task, so real savings will run smaller than the 90 percent sticker — the same dynamic we saw when Opus's token counts jumped after its own tokenizer change. The Sonnet discount builds on Anthropic ships Sonnet 5.5 as cyber safeguards move down a tier.
Liquid AI open-sourced the d1 line: two small decision models, d1-3B and d1-omni-600M. Decision models are a different animal from Liquid's generative LFM family — they produce an answer in a single forward pass with no output tokens, which is why the latency numbers look absurd: 8 milliseconds on an RTX 4090, 16 ms on a Jetson Thor. The 3.12-billion-parameter d1-3B posts a 48.57 on the Decision Index 0.2.1 benchmark, the best score below 10 billion parameters and ahead of Decider 35B-A3B at 47.11 — though those are Liquid's own numbers, not an independent run. NVIDIA co-engineered the Jetson versions and llama.cpp support shipped day one, but note the license is Liquid's own LFM Open License rather than an OSI-approved one, so "open weights" comes with strings.
Former White House AI adviser Sriram Krishnan is reportedly raising a $500 million fund for US AI startups. The Information, citing three people familiar with the matter, reports the fund targets growth- and later-stage companies with national-security relevance for the US and allied nations; Axios independently confirms the roughly $500 million target and adds that there's no fund name yet and the team is still being assembled. Krishnan, an ex-a16z partner, left his White House role in June. Nothing is closed and he hasn't commented — but the pitch tells you where post-Washington AI money expects to go.
What to watch: whether OpenAI or Google answer Anthropic's cuts within days, and whether Krishnan lands anchor LPs before year end.
Is the Haiku 5.5 price cut a real efficiency gain or a tokenizer accounting trick that evens out in production? Tell us in the comments.



