What is adversarial machine learning?

By Identra · Updated

Adversarial machine learning is the study of deliberate attacks on machine learning systems and the defenses against them. The attacks aim to change a model's predictions, corrupt its training, copy its behavior, or pull sensitive information out of it.

How do adversarial machine learning attacks work?

An attacker goes after what the model learned, how it reacts to input, or what its output gives away. They might slip examples into training data. They might send crafted queries to a prediction API. Against an LLM app they may only need to plant text somewhere the app will read it.

An adversarial example is input built on purpose to get the wrong answer. Picture a malware sample tweaked so it still runs but the classifier now calls it clean. A model being wrong by itself isn't an attack. Someone has to be steering it.

The NIST adversarial machine learning taxonomy sorts attacks by the attacker's goals, capabilities and knowledge, and by the lifecycle stage they hit. For an enterprise assessment, start smaller. What asset is at risk, and what access would the attacker need to reach it?

What are the main types of adversarial ML attacks?

Four families come up again and again. They overlap. A backdoor, for example, gets planted through poisoning and fires later when a trigger input shows up.

Model extraction steals what a model does. Privacy attacks go after the data or the people behind it. Leaking a chatbot's system prompt is a different thing again. It exposes the app's instructions and leaves the model alone.

  • Evasion

    What the attacker wants
    A wrong prediction from a deployed model
    Example
    Tweak a fraudulent card transaction so the fraud model scores it as normal
  • Poisoning

    What the attacker wants
    Corrupted learning or a hidden behavior
    Example
    Slip mislabeled samples into the next retraining batch
  • Extraction

    What the attacker wants
    A copy of the model's functionality
    Example
    Query a prediction API at scale and train a lookalike on the answers
  • Privacy inference

    What the attacker wants
    Sensitive facts about the training data
    Example
    Work out whether a specific patient's record was in the training set

What access does an attacker need?

White-box means the attacker can see internals like weights or gradients. Black-box means they send queries and read responses. Real attackers usually sit somewhere in between, so write down exactly what they can see and change.

AI model poisoning needs a hand in the training data, the training process or the model files. Evasion, extraction and privacy attacks can work with query access alone.

Then get specific. Does the attacker need an authenticated account? Does the API return confidence scores or just a label? How many queries before a rate limit kicks in? A lab demo with full internal access says little about what an outsider can do through a locked-down production endpoint.

How does adversarial ML apply to LLMs and AI agents?

LLMs inherit poisoning, extraction and privacy risks. They also take instructions in plain language from anyone whose text reaches the context window. That's prompt injection. Jailbreaking is its cousin, aimed at talking the model out of its safety rules.

Keep training-time and runtime problems apart. A poisoned SharePoint page that Microsoft 365 Copilot pulls into an answer can change the output without touching a single weight. The investigation looks at the document, the retrieval path and what the assistant did next. The training pipeline may be perfectly fine.

Say a support agent reads a ticket telling it to look up other customers and email their records out. How bad that gets depends on what the agent can read and where it can send. Agentic AI security has to deal with the manipulated reasoning and with the actions that reasoning can trigger.

How do you defend against adversarial ML?

Tie each control to an attack path. If you train models, you need control over data and model changes. If you only call the OpenAI or Anthropic API, you still own your API keys, the context you send, the data you connect and what the app is allowed to do.

Apply least privilege to the identity that actually executes each action. The support agent fetches records for the ticket in front of it and nothing else. Enforce that in the data service. Good model behavior is not a control.

  • Track where datasets and models came from. Restrict who can change them and keep known good versions to roll back to.
  • Authenticate API clients, set usage limits and look into strange query patterns. This slows extraction down. It doesn't make it impossible.
  • Keep sensitive data out of training sets where you can, and test privacy protections against what an attacker would actually try to infer.
  • Treat retrieved documents and tool responses as untrusted. Authorization decisions stay outside the model.
  • Require explicit approval for external transfers and destructive changes, and show the reviewer the real target.
  • Keep access-controlled records of inputs, model versions, tool requests and outcomes so an incident can be reconstructed.

How do you test whether the defenses work?

Point AI red teaming at the deployed workflow with realistic attacker access. Use the real retrieval sources, permissions, rate limits and tools. Score model behavior and downstream outcomes separately. A manipulated answer is one finding. An unauthorized data transfer is a different and worse one.

For the support-agent case, drop a test ticket with hostile instructions into staging. Did the agent try the lookup? Did the data service refuse it? Did the outbound email stop for approval? Use synthetic records and a mailbox you control.

Adversarial training helps against the attacks it was trained on. New models, data sources, tools or permissions mean running the tests again, and every open failure gets a named owner.

How Identra thinks about it

Identra covers the enterprise side, where AI is in use. AI agent tool calls are checked against policy, and destructive shell commands are denied when policy is set to block. Prompts are checked on the device before they are sent. Connected AI apps and OAuth grants are sorted by what they can reach, and every AI agent run is recorded with user, device and outcome for investigation.

Go deeper: AI security, built on identity

Frequently asked questions

Is adversarial machine learning the same as using AI for cyberattacks?

No. Adversarial ML is about attacking machine learning systems. Using an LLM to write phishing emails is AI as an attack tool. One operation can involve both.

What is the difference between evasion and poisoning?

Evasion feeds crafted input to a deployed model to get a wrong result. Poisoning tampers with training data, the training process or model files to change how the model behaves later.

Does a privacy attack recover exact training records?

Not necessarily. Membership inference estimates whether a record was in training. Model inversion infers attributes or reconstructs representative data, which isn't the same as recovering the original record.

Does prompt injection require access to model weights?

No. The attacker only needs to control some text the app reads, like an email, a document or a web page. The model itself stays untouched.

Can enterprises using hosted models ignore adversarial ML?

No. The provider handles parts of model security. You still own the app's permissions, the data you connect and the actions you allow, so test those.

Related terms

Keep exploring · AI threats and attacks