AI 101 — What is overfitting?

Share
AI 101 — What is overfitting?

Overfitting is when a model learns its training data too well — it gets the examples it has already seen almost perfect, then performs worse on anything new. Since the entire job of a model is to handle cases nobody showed it, overfitting is the most common way an AI system looks brilliant in a demo and disappoints in production.

The term shows up in almost every model release, fine-tuning guide and benchmark write-up, and it is rarely explained. The short version: a model is only worth anything if what it learned generalises, and the failure to generalise is called overfitting. Its opposite is underfitting — a model so weak or so briefly trained that it misses the pattern in the data it already has.

Why it matters right now

Three places it bites in 2026.

Fine-tuning on small data. When a team adapts a model to its own material — support tickets, legal contracts, a codebase — the usual dataset is a few thousand examples, sometimes a few hundred. That is exactly the regime where a model can memorise its examples rather than learn from them. The fine-tuning route therefore carries a measurement obligation, which is what our guide to deciding if a model needs fine-tuning puts ahead of touching the weights: build the held-out test set first, or you cannot tell learning from memorising.

Big models were once badly undertrained. In 2022 a DeepMind team trained over 400 language models, from 70 million to more than 16 billion parameters, on 5 billion to 500 billion tokens of text, and found that for a fixed compute budget, model size and training data should grow together — double the parameters, and you should double the data too. They then built Chinchilla, a 70-billion-parameter model trained on four times more text than Gopher, and it beat Gopher (280 billion parameters), GPT-3 (175 billion) and Megatron-Turing NLG (530 billion). The bigger models did not lose on brains. They had been given too little data for their size — which is a form of underfitting, and it was the default setting across the industry until someone measured it.

Benchmark scores that don't survive contact. The mirror image of overfitting to your training set is overfitting to your test set. When a test's questions leak into the training data, scores inflate, and a whole literature now exists to detect it. It is why the site keeps returning to rigged evaluations: models cheating on cyber benchmarks and Google's pilot of double-blind evaluation are the same problem wearing different clothes.

Two students focused on an exam in a classroom setting during daylight.

The one-paragraph mental model

Training works by showing a model examples and nudging its internal numbers until its answers match. If you only ever measure that matching on the examples being used to train, the number always improves — right up to the point where the model is reproducing its inputs instead of extracting the rule behind them. The fix is structural: hold a slice of data back, never let the model train on it, and watch whether the score on that slice improves too. Two curves, then. When both fall together, the model is learning. When training error keeps falling while held-out error flattens and turns upward, it has started memorising — generally visible long before the training score looks suspicious. The remedies are unglamorous and known: more data, stop earlier, keep the model smaller, or add regularisation that penalises needless complexity. There is genuinely no way to tell memorising from learning by looking at the training results alone, which is why a held-out set is not optional.

The exam analogy

Picture a student who studies for a test by memorising last year's paper, question by question, answer by answer. Sheet one, the past paper, and the score is near perfect. Sheet two, the same topics reworded, and it collapses — because what was learned was the exact phrasing, not the subject.

That is overfitting, almost precisely. A well-prepared student does something different: they extract the rules, check themselves on questions they have not seen, and revise when the unfamiliar ones go badly. That checking step is a validation set.

Google's machine-learning course uses a matching example for models. Imagine separating healthy trees from sick ones on a map of a forest. Draw a single rough boundary and you misclassify a few trees but probably get new trees roughly right. Draw an elaborate squiggle that snakes around every individual tree to classify the training forest perfectly, and it will do badly the moment a new plot of trees arrives. Complex enough to fit everything is usually complex enough to fit the noise.

Common misconceptions

"An overfit model is too big." Sometimes, but size alone predicts nothing. Modern networks are routinely far larger than their training data, and the classic U-shaped trade-off between too simple and too complex breaks down past a certain point — beyond the size where a model can fit its data exactly, error can start falling again instead of rising. Big models are less prone to this, not more, which is why "just use a smaller model" is not the fix.

"Overfitting means the model was trained too long." Training duration is one cause, not the definition. Too little data, too much duplication, a dataset that is all one flavour, or training on the test set all produce the same symptom. The definition is about the gap between seen and unseen, not the clock.

"Low training loss means a good model." It means the model fits what it was shown. That is why a number in a launch post that only describes training performance tells you nothing about whether the system works — and why a benchmark a model was accidentally trained on is a broken measurement rather than strong evidence.

"Overfitting only matters to people who train models." Anyone who has adapted a model to their own data is training one. The mitigation is not a technique you need to implement; it is a discipline of keeping a slice of examples the system has never seen, and re-running them every time you change anything.

Where to learn more

Google's machine-learning crash course has the clearest plain-English treatment of overfitting and underfitting. The scikit-learn library's own example plots both failures against a good fit in three panels, which is worth more than a paragraph of explanation. The double-descent paper is where the classical rules are renegotiated for modern networks; the grokking paper is the best reminder that generalisation in deep learning can arrive long after the training curve looks finished — and the Chinchilla paper is the clean example of a whole industry fixing an underfitting problem it had mistaken for a capability ceiling.

Related reading: What is backpropagation? is the mechanism that drives the training curve down in the first place · What is fine-tuning? is where overfitting does the most practical damage · and What is an AI benchmark? covers the test set leaking into training data.

Have you shipped something that worked perfectly in testing and failed in front of a real user? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t