AI 101 — What is a foundation model?

Share
AI 101 — What is a foundation model?

A foundation model is one AI model trained on a huge, general pile of data and then adapted to many different jobs — a single base that chatbots, coding assistants, search features, and scientific tools are all built on top of. Instead of building a separate model for every task, companies build (or license) one foundation model and specialize it.

Why it matters right now. The AI industry stopped shipping individual products and started shipping model families. This week alone, Anthropic released Claude Haiku 5.5 as a cheaper member of its Claude lineup, and OpenAI's GPT-6 safety report counted fewer refusals but more regressions — both read as trade-offs inside one family, not two unrelated launches. The economics push the same way: pretraining a frontier model takes thousands of chips and months of work, so almost everyone builds on someone else's foundation rather than their own. And regulators have landed on the same unit of analysis — Europe's rules attach legal obligations specifically to general-purpose AI models, the regulator's name for foundation models. When one artifact carries the product roadmap, the safety debate, and the law, it is worth knowing exactly what it is.

The mental model. Think of it as two stages. First comes pretraining: the model reads an enormous, internet-scale corpus and learns to predict patterns — next word, next pixel, next amino acid — without any particular task in mind. Nothing about "answering support emails" or "finding drug candidates" is in this stage; the model is acquiring raw competence, the way a medical student learns anatomy before picking a specialty. Second comes adaptation: the same base is pointed at a real job through prompting, retrieval (we covered the mechanics in What is RAG?), or fine-tuning on domain examples. The expensive, once-in-a-generation work happens in stage one; nearly everything a product team does happens in stage two.

The analogy. Picture a restaurant kitchen. A foundation model is a cook who trained for years across every cuisine — technique, flavor pairings, knife work — but has never seen your menu. When hired, the cook doesn't go back to culinary school; you hand over the house recipes and a few service nights of feedback, and the repertoire comes together quickly. That's why one kitchen can swap cooks without rebuilding the restaurant, and why the expensive part (the years of training) is shared while the cheap part (learning the menu) is per-restaurant. The menu is the product; the cook is the foundation.

Common misconceptions. First: "a foundation model is a chatbot." A chatbot is one application draped over the base — and plenty of foundation models never chat at all. Embedding models behind semantic search, image generators, and models trained for science are foundation models too.

Second: "foundation model means large language model." Language models are the most visible kind, but the defining feature is the recipe — broad pretraining plus adaptability — not the modality or size. A large language model is the most common passenger; it is not the car.

Third: "foundation means finished." The Stanford report that popularized the term in August 2021 chose it deliberately to signal "critically central yet incomplete" — the models underpin everything and still fail at reliability, planning, and up-to-date knowledge without help. They are a foundation, not a building.

Fourth: "foundation model equals open weights." Unrelated axes: open-weights models describes how a model is released; foundation describes how it was built. Open-weight releases (Llama, DeepSeek) are foundation models; most closed ones are too.

Where to learn more. Start with Stanford's On the Opportunities and Risks of Foundation Models — the report that named the category; its introduction alone is readable in twenty minutes. For the moment the recipe went mainstream, OpenAI's GPT-3 paper (175 billion parameters, no fine-tuning needed for new tasks) shows stage one working in the wild. And for the money question — build on someone else's foundation or train your own — our explainer What is fine-tuning? covers the adaptation side most teams actually touch.

Related reading: What is a large language model? · What are open-weight models? · What is AI inference?

When a lab ships a new model family, what decides your take — raw capability, or what it costs to run? Tell us in the comments.

Read more

Akhetonics says its all-optical CPU reaches a customer in 2026

Akhetonics says its all-optical CPU reaches a customer in 2026

Light-based computing keeps promising more than it delivers — but one Munich startup has just put a date on its bet, and the interview laying it out is doing the rounds on Hacker News this week. Akhetonics says it will deploy its first commercial machine with a major customer by the end of 2026, with several more planned for 2027. The company, founded by Michael Kissner and Leonardo Del Bino, is building a computer where data enters as light, is switched as light, and circulates through memory

The Week in AI — October 5–11, 2026

The Week in AI — October 5–11, 2026

Every big claim this week turned out to rest on fine print more interesting than the headline: revenue only the company reporting it can define, safety tests sandboxed while the product keeps the web, and a Pentagon phase-out nobody would confirm until reporters kept asking. The week's top 5 1. OpenAI's revenue was $20 billion below the numbers everyone quoted — and the gap was definitional. The Financial Times reported Thursday that OpenAI's annualized revenue runs roughly $20 billion unde

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Drone strike shuts a third Yandex data center, taking YandexGPT offline

Russia's largest tech company is learning what the AI era's infrastructure war looks like from the receiving end — three data centers in four days, and with them much of the cloud layer Russian businesses run on. A Ukrainian drone strike knocked out Yandex's data center in Vladimir early Sunday morning, the third of the company's facilities hit since October 8. The site — reported at roughly 50 MW and designed for about 2,880 server racks — stopped operating completely after the attack, Yandex

Agent teams cost up to 5x more, barely score higher

Agent teams cost up to 5x more, barely score higher

The multi-agent hype train hit a benchmark this weekend — and the grid and the trucking regulators had quiet weeks of their own. Vals AI put agent teams head-to-head with single agents on its Vibe Code Bench, and the teams cost between 1.8 and 5.1 times more for almost no extra quality. The evals company ran GPT-6 Sol and Claude Opus 5.5 solo and in teams across 50 apps at two reasoning efforts; out of four comparisons, only one was statistically significant — Sol at medium effort, where the t