AI 101 — What is a foundation model?

A foundation model is one AI model trained on a huge, general pile of data and then adapted to many different jobs — a single base that chatbots, coding assistants, search features, and scientific tools are all built on top of. Instead of building a separate model for every task, companies build (or license) one foundation model and specialize it.
Why it matters right now. The AI industry stopped shipping individual products and started shipping model families. This week alone, Anthropic released Claude Haiku 5.5 as a cheaper member of its Claude lineup, and OpenAI's GPT-6 safety report counted fewer refusals but more regressions — both read as trade-offs inside one family, not two unrelated launches. The economics push the same way: pretraining a frontier model takes thousands of chips and months of work, so almost everyone builds on someone else's foundation rather than their own. And regulators have landed on the same unit of analysis — Europe's rules attach legal obligations specifically to general-purpose AI models, the regulator's name for foundation models. When one artifact carries the product roadmap, the safety debate, and the law, it is worth knowing exactly what it is.
The mental model. Think of it as two stages. First comes pretraining: the model reads an enormous, internet-scale corpus and learns to predict patterns — next word, next pixel, next amino acid — without any particular task in mind. Nothing about "answering support emails" or "finding drug candidates" is in this stage; the model is acquiring raw competence, the way a medical student learns anatomy before picking a specialty. Second comes adaptation: the same base is pointed at a real job through prompting, retrieval (we covered the mechanics in What is RAG?), or fine-tuning on domain examples. The expensive, once-in-a-generation work happens in stage one; nearly everything a product team does happens in stage two.
The analogy. Picture a restaurant kitchen. A foundation model is a cook who trained for years across every cuisine — technique, flavor pairings, knife work — but has never seen your menu. When hired, the cook doesn't go back to culinary school; you hand over the house recipes and a few service nights of feedback, and the repertoire comes together quickly. That's why one kitchen can swap cooks without rebuilding the restaurant, and why the expensive part (the years of training) is shared while the cheap part (learning the menu) is per-restaurant. The menu is the product; the cook is the foundation.
Common misconceptions. First: "a foundation model is a chatbot." A chatbot is one application draped over the base — and plenty of foundation models never chat at all. Embedding models behind semantic search, image generators, and models trained for science are foundation models too.
Second: "foundation model means large language model." Language models are the most visible kind, but the defining feature is the recipe — broad pretraining plus adaptability — not the modality or size. A large language model is the most common passenger; it is not the car.
Third: "foundation means finished." The Stanford report that popularized the term in August 2021 chose it deliberately to signal "critically central yet incomplete" — the models underpin everything and still fail at reliability, planning, and up-to-date knowledge without help. They are a foundation, not a building.
Fourth: "foundation model equals open weights." Unrelated axes: open-weights models describes how a model is released; foundation describes how it was built. Open-weight releases (Llama, DeepSeek) are foundation models; most closed ones are too.
Where to learn more. Start with Stanford's On the Opportunities and Risks of Foundation Models — the report that named the category; its introduction alone is readable in twenty minutes. For the moment the recipe went mainstream, OpenAI's GPT-3 paper (175 billion parameters, no fine-tuning needed for new tasks) shows stage one working in the wild. And for the money question — build on someone else's foundation or train your own — our explainer What is fine-tuning? covers the adaptation side most teams actually touch.
Related reading: What is a large language model? · What are open-weight models? · What is AI inference?
When a lab ships a new model family, what decides your take — raw capability, or what it costs to run? Tell us in the comments.



