What is data and model poisoning?
By Identra · Updated
Data and model poisoning is deliberate tampering with training data, fine-tuning examples, model weights or model updates so that a model behaves the way an attacker wants. The result can be lower accuracy, a steered decision, or a backdoor that only fires on a specific trigger.
How does data and model poisoning work?
The attacker needs a hand in what the model learns from. That could be a public dataset scraped into a training run, labels from an outside contractor, thumbs-up feedback that gets folded into fine-tuning, or a LoRA adapter pulled from Hugging Face. With write access to the model file itself, they can skip the data and edit the weights.
Some attacks just make the model worse. Others are targeted. A backdoored spam filter can pass every test you run and still wave through any email that contains one odd phrase. NIST lays out these categories in its adversarial machine learning taxonomy.
Poisoned examples don't always take. Whether they stick depends on the training process, how much clean data surrounds them and the task. Before you call a bad output poisoning, find out what an attacker could actually have changed.
How is model poisoning different from RAG poisoning or prompt injection?
It comes down to where the attacker touches the system. Training poisoning changes what the model learned. Retrieval poisoning changes the documents handed to the model at answer time, and the weights stay as they were. Prompt injection puts instructions straight into the model's input.
The lines blur. A wiki page with hidden instructions is retrieval poisoning and indirect prompt injection at once. A page with a wrong bank account number and no instructions at all can still mislead the answer. So RAG security has to ask whether a source is true, as well as whether it's hostile.
Training data poisoning
- What the attacker changes
- Examples or labels the model learns from
- How you recover
- Clean the data, then roll back or retrain
Direct model poisoning
- What the attacker changes
- Weights, adapters or model updates
- How you recover
- Restore a verified artifact and find out how it was changed
Retrieval poisoning
- What the attacker changes
- Documents fed in as context
- How you recover
- Fix the source, then purge index and cache copies
Prompt injection
- What the attacker changes
- Instructions inside the model's input
- How you recover
- Close the input path and limit what the model can trigger
What could poisoning look like at a company?
Say a support team fine-tunes a model on resolved Zendesk tickets every quarter. An attacker files a batch of tickets with a plausible fix that tells customers to turn off MFA, then upvotes those answers. Nobody reviews the training set. The next release starts recommending it.
Tracing that means getting from the bad answer to a model version, from the version to a dataset snapshot, and from the snapshot to whoever approved it. Without those links you're guessing.
A retrieval case plays out differently. Someone edits the vendor payment page in Confluence, and the finance assistant quotes the new bank details with a neat citation. The citation shows where the text came from. It says nothing about whether the text is legitimate. Fix the page, then re-index and clear cached copies. In the fine-tuning case, deleting the bad tickets from the dataset wouldn't undo what the model already learned.
How do you protect training data and model releases?
Treat the right to promote data into a training set as a privileged action. Apply least privilege to the people and service accounts that collect examples, approve datasets, run training jobs and push models to production.
- Version every dataset and record who approved it.
- Hold user feedback and outside contributions in quarantine until someone reviews them.
- Review changes to labeling rules and ingestion scripts the way you review model code.
- Keep the evaluation set apart from training data, with its own write permissions.
- Before release, test the decisions that matter most and the inputs an attacker would pick.
- Keep known-good datasets and checkpoints so a rollback never depends on rebuilding from suspect inputs.
What should you check in third-party models?
Downloaded models, adapters and datasets belong in AI supply chain security reviews. Know the publisher, where the file came from, which version you approved and how updates arrive. A hash tells you the file matches a reference copy. A signature can tell you who published it. Neither tells you anything about the training data.
For a model you call over an API, the questions go to the supplier. How do they vet training inputs? How will they tell you when the model behind an endpoint changes in a way that matters?
AI red teaming can probe for targeted failures that normal QA misses. A clean result covers what you tested and nothing more. Keep separate approval checks on payments and account changes regardless.
What should you do if you suspect poisoning?
Pull the affected workflow out of service and save the evidence. That means the model version, dataset references, retrieved documents, prompts, outputs and change history. Run the same inputs against a known-good release. One wrong answer doesn't prove tampering.
Recovery depends on which layer was hit. Poisoned training data or weights mean rolling back to a trusted checkpoint or retraining on reviewed data. Poisoned retrieval means fixing the source and purging stale copies from the index and cache.
Then follow the damage downstream through AI incident response. Which tickets, refunds or config changes relied on the bad output? Their owners need to know. Close whatever access gap let the bad data in before the pipeline runs again.
How Identra thinks about it
Identra limits what a compromised model or AI tool can do on a work laptop. On macOS and Windows it finds the AI clients, coding agents, MCP servers, plugins and packages in use. AI agent tool calls are checked against policy, destructive shell commands can be denied when policy is set to block, and every agent run is recorded with user, device, AI client and outcome for the investigation.
Go deeper: AI security, built on identity
Frequently asked questions
Is every incorrect AI answer evidence of poisoning?
No. Gaps in the data, retrieval misses and plain model limits cause most wrong answers. Poisoning means someone did it on purpose, so look for source changes and reproducible behavior first.
Can poisoning affect a company that never trains models?
Yes. You can inherit it through a third-party model or adapter. Your retrieval system can also pick up poisoned documents, though that leaves the model itself unchanged.
What is a model backdoor?
Hidden behavior that a specific trigger switches on. The model acts normally otherwise, so ordinary quality testing can miss it.
Does deleting poisoned training data fix the model?
No. The model has already learned from it. You need to restore a trusted model or retrain, then run targeted evaluations.
Can a signed model still be poisoned?
Yes. A valid signature shows who signed the file and that it hasn't changed since. It says nothing about whether the training data was clean.
