How to — estimate what an AI feature will cost before you ship it

Share
How to — estimate what an AI feature will cost before you ship it

You want one number you can defend in a planning meeting: what each completed job on your feature costs, and what a month of it costs at launch volume. This is a measurement job, not a guess — a morning with your real prompts beats any pricing page.

1. Run 20 real requests and keep the usage numbers. Not demo prompts — your actual system prompt, tool definitions, the documents you stuff into context, and a typical user message. Every serious API returns input and output token counts with each response; log them, then take the median (averages get dragged by one monster request). If you're building on a platform that doesn't report usage, run the same twenty prompts through the provider's own console and read the usage report there. What you're collecting is the size of "one call" in the only unit the bill cares about. If the anatomy of those units is new to you, AI 101 — What is a token in AI? is the two-minute version.

2. Price that call on three meters, not one. Input tokens, output tokens, and — on any model with reasoning or an effort dial — the thinking tokens, which are billed as output even though your user never sees them (see AI 101 — What are reasoning tokens?). Output rates are typically a multiple of input, so a chatty model punishes you twice: it writes more, and each written token costs more. Then check the fine print for the long-context cliff — Claude Haiku 5.5, for example, charges one rate up to 100,000 tokens of context and five times that above it. If your feature routinely passes six figures of context, you're not pricing the sticker, you're pricing the tier above it.

3. Read the discount lines like a lawyer. Caching discounts repeated context — real only if your prompts actually repeat: a fixed system prompt qualifies, a fresh agentic transcript may not. Batch or off-peak submission is cheaper, but only for work that can wait hours for its answer. Apply only the discounts your traffic genuinely qualifies for. An estimate with every discount stacked on it is a number for a different product than yours.

4. Multiply by the loop, not the call. A chat box makes one call per turn; an agent doing a job makes ten, thirty, sometimes more — and each trip re-reads its context. Run your feature through one complete job on real data, count the model calls in the trace, and multiply. Add headroom for retries and failed steps. Then re-run your token counts against the exact model version you'll ship on: tokenizer changes between versions move the same prompt's token count by double-digit percentages, and that's an invisible price rise on every request.

5. Stress the estimate on purpose. Twist each dial to its realistic worst case: default effort but occasionally high, outputs running to the cap instead of the median, traffic a few multiples of your launch guess. For scale, Simon Willison's hands-on test of Haiku 5.5 found the same picture-of-a-pelican prompt costing 0.0936 cents at low effort and 3.3826 cents at max — about 36 times more for one prompt, on one sticker price. You don't budget for the max, but you must know what it costs before someone flips it.

6. Write the budget as a ceiling, then alarm on it. One line each: this feature may cost X per completed job, and Y per month at N jobs a day. Attach a spend alert to the API account before launch, not after. Cost work fails silently — nobody notices a bill until it's quarterly.

Don't do this: don't price from the launch headline. "Up to 90% cheaper" and "$0.10 per million tokens" are marketing meters; your invoice is per unit of work. As we wrote in Deep Dive — Haiku 5.5's fine print: same price, three times the bill, Haiku 5.5 and GPT-6 Luna quote identical per-token prices at their top effort settings, yet measured per-task cost ran $0.21 against $0.07 — three times the bill on the same sticker. Quote from your own measurement, never from someone else's chart.

How you'll know it worked: three or four weeks after launch, the actual invoice lands within about 30 percent of your estimate, and the per-job cost you tracked matches the invoice divided by jobs. If it's off by more, exactly one meter lied — usually output volume from a chattier-than-expected model, or a loop longer than the one you traced. Fix that meter and re-run the arithmetic; the method holds.

What does one completed job on your AI feature actually cost — and did measuring it change what you ship? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t