AI 101 — What is data poisoning?

Data poisoning is slipping malicious examples into the data an AI system learns from, so the model quietly picks up a lesson the attacker chose — while behaving completely normally on everything else. The finished model passes its tests; the bad habit is baked into its weights, waiting for the right trigger.
Why it matters right now
Poisoning used to sound like a theoretical worry aimed at giant pretraining runs. Three developments in the last few years moved it into everyday territory.
First, the training corpora are scraped from the open web, and the open web is mutable. A team led by Nicholas Carlini showed in 2023 that an attacker could poison 0.01% of two popular image-text datasets — LAION-400M and COYO-700M — for about $60, by exploiting the fact that a web page can show a dataset's annotator one thing and everyone who downloads it later something else. A second route: briefly editing crowd-sourced pages like Wikipedia in the window between two periodic snapshots.
Second, the models that write code are trained on unvetted public code. Hojjat Aghakhani and colleagues demonstrated attacks — COVERT and TrojanPuzzle — that hide poison inside something as boring as a docstring, so that static scanners looking for the malicious payload never find it. The poisoned model then volunteers insecure code during ordinary, harmless suggestions.
Third, and newest: AI agents now learn after deployment. A 2026 paper from Emory University showed that an attacker doesn't need dataset access at all — submitting ordinary-looking tasks can steer an agent's own execution history until the system consolidates the pattern into a reusable skill containing a hidden trigger. The attacker's only foothold is the ability to type.
It is no longer a niche concern: the OWASP GenAI Security Project lists Data and Model Poisoning as one of its standing Top 10 risks for LLM applications.
The mental model
Training works by example: show a model enough examples and it infers the pattern. Poisoning exploits that directly. An attacker has two playbooks. The blunt one is degradation — feed in enough junk and the model gets measurably worse at everything; controlled pretraining runs reported in September 2026 found the model's clean-data error climbing steadily as the poison rate grows, even when no backdoor is installed. The precise one is the backdoor: plant a small number of examples that all pair an otherwise-invisible condition — a specific word, a file pattern, a phrase — with a behavior the attacker wants, and the model learns "when this appears, do that." On every other input it performs normally, which is exactly what makes the attack hard to catch. Crucially, the malicious examples are crafted to look legitimate; work on recommender systems as far back as 2016 showed attackers generating fake entries that mimic normal user behavior precisely to avoid detection.
The analogy
Think of a reservoir that supplies a town's drinking water. A rival pours a trace contaminant into it — so little that every routine safety test comes back clean, day after day. Most residents never notice anything. But anyone who drinks a glass after eating a particular food feels ill, every time. The supply looks and tastes normal; the harm is conditional on a trigger no inspector is testing for. That is a poisoned model: normal in the demo, dangerous on command.
Common misconceptions
"A poisoned model would act broken everywhere." That is the failure mode attackers avoid on purpose. Backdoor attacks are explicitly evaluated on preserving normal performance — the 2026 skill-poisoning paper reports implants that kept clean-task utility intact, matching or beating straightforward injection. The model you benchmark is healthy; only the trigger reveals the lesson.
"Poisoning means hacking into a dataset." Sometimes the data simply changes hands: the $60 web-scale attack required no break-in, only a web page whose content differed between two visits. And the idea predates chatbots — poisoning research targeted recommender systems a decade before large language models existed.
"Poisoning, jailbreaking, and prompt injection are the same scare with different names." They differ in timing and in how hard they are to remove. What is prompt injection? is hostile instructions arriving through data the model reads right now — fix the document and the attack is gone. A jailbreak talks the model past its rules during a conversation, and its effect ends with the conversation. Poisoning happens while the model learns, so the malicious behavior lives in the weights themselves — you cannot delete a poisoned training example the way you delete a poisoned file, which is why labs rehearse these scenarios through What is AI red-teaming? before outsiders do.
"Fine-tuning or safety training will scrub it out." Not reliably. Backdoors surviving further training is common enough that it is a standard subject of the poisoning research literature — one reason models that keep learning from new inputs, as covered in What is continual learning?, carry the exposure forward rather than closing it.
Related reading: What is prompt injection? · What is AI red-teaming? · What is continual learning?
If you cannot inspect the data a model learned from, can you trust what it learned? Tell us in the comments.



