AI 101 — What is AI inference?

Share
AI 101 — What is AI inference?

Every answer an AI gives you is inference: running a trained model on new input to produce an output. Training is how a model learns; inference is how it works. If training is teaching, inference is doing.

Why it matters right now. Training makes the headlines — a new model, a bigger cluster, a record run — but training happens once per model, while inference happens every time anyone asks a question. That asymmetry is why the money has shifted: Google Cloud describes inference as the phase "where AI delivers business value," and this week's news — Amazon pledging $1 billion to host data-center towns, an AI unit reinventing itself as an inference cloud — is infrastructure being built for that steady, always-on demand, not the one-off act of training. When people talk about "inference economics," this is what they mean: the cost of answers, multiplied by billions of answers.

The mental model. A trained model is a finished instrument. Its weights — the billions of numbers learned during training — are frozen at release; nothing new is learned when you type a message. When your request arrives, the model reads your prompt in one pass, then writes its reply one token at a time, each new token predicted from everything before it. Two phases, both called inference: reading the question (engineers call it prefill) and writing the answer (decode). Your whole chat window is a sequence of these cycles, repeated as fast as the hardware allows.

Contemporary computer with black screen placed on stand near row of server steel racks in data center

The analogy. Think of a driver who spent years in driving school. That was training — expensive, slow, done once. Now she drives a taxi route eight hours a day. Every trip is inference: same knowledge, new street, instant judgment. Each individual trip costs a tiny fraction of what the school cost, but the taxi company's entire budget is trips, because she does thousands of them. The school made her a driver; the trips are the business.

Common misconceptions. First: "the model is learning from our conversation." It isn't. Inference runs the trained weights as-is — which is why a chatbot can be confidently wrong and why your corrections don't stick (when a vendor does let a model learn from feedback, that's a separate step, usually fine-tuning). Second: "inference is the cheap part." Per request, yes — Google Cloud notes each prediction is far less computationally demanding than a training run. But training a frontier model happens a handful of times a year, while inference runs continuously at global scale, so it now dominates how AI compute is actually spent. Third: "inference" doesn't mean the model is reasoning or concluding anything deep — the word is borrowed from statistics, where it simply means drawing an output from a model given data.

Why inference has its own industry. Because it runs constantly and users watch the clock, engineers optimize it differently from training. Requests from many users are batched onto the same hardware; a model can be quantized — its numbers stored with less precision — to fit more answers per second; and repeat portions of your prompt can be cached so the model doesn't re-read your whole document every message (as we covered in What is prompt caching?). Even the economics of thinking changed: reasoning models spend extra inference — tokens generated before the visible reply — to work through hard problems, which is why turning "thinking" up makes the same question cost more. Every token you see is produced one at a time by this machinery, one predicted from what came before.

Related reading: What is a large language model? · What is a token in AI? · What is model quantization?

When a chatbot gets something basic wrong, do you blame the model or the prompt? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t