What is AI red teaming?

By Identra · Updated

AI red teaming is authorized adversarial testing that looks for ways an AI system can break its security or safety requirements. It covers the model and the application around it, to show whether an attacker can expose protected data, get around restrictions or trigger unauthorized actions.

What does AI red teaming test?

Whatever the deployment can get wrong. A public chatbot gets tested for harmful content and privacy slips. An internal assistant with access to Salesforce, a SharePoint index and an email tool needs much more. Can someone use it to read records they should not see? Can they make it send something?

Write the forbidden outcomes down first. "Shows customer A's invoices to customer B." "Sends an email without approval." NIST's Generative AI Profile recommends adversarial testing to find failure modes nobody expected.

Say which risks are in scope. A two-day prompt injection exercise is useful. It is not a safety review.

What does an AI red team test look like?

Say a support team uses an assistant that reads new tickets, looks up the customer and drafts a reply. With sign-off, a tester files a synthetic ticket. Hidden in it is an instruction to pull a different customer's records and email them to an outside address the tester controls. That is indirect prompt injection. The attack arrives as ordinary work.

Follow it all the way through. Did the assistant treat the ticket text as a command? Did it ask for the other customer? Did the data service hand the records over? Did the email tool send?

Each answer is a separate finding. If the assistant tried the lookup and the data service refused, you have a model failure and a control that held. If the synthetic records landed in the tester's inbox, you have a disclosure path and an owner to call.

How is it different from penetration testing or jailbreaking?

The three overlap. Pen testing still matters for authentication, broken authorization and exposed secrets. Jailbreaking is narrower than either. A jailbreak can get a forbidden answer out of a model without touching company data. And an assistant can leak a restricted file through a broken permission check with no jailbreak involved.

  • Penetration testing

    Looks for
    Exploitable app and infrastructure flaws
    Example finding
    A user pulls another user's records through the API
  • Jailbreak testing

    Looks for
    Ways around the model's behavior rules
    Example finding
    A crafted prompt produces banned content
  • AI red teaming

    Looks for
    Failures across model behavior, data access and actions
    Example finding
    A poisoned document triggers an unauthorized export

How do you run a useful exercise?

Give testers realistic access. Someone who can file a ticket, edit a shared doc or log in as a regular employee. Hand them admin and you learn nothing about the attacker you actually worry about.

Agree on targets, stop conditions and cleanup before anyone starts. Use fake customer records and a destination you own, so a successful AI data leakage test leaks nothing real. Keep the assistant's production permissions. A test account with broader rights produces findings that will not reproduce.

Mix scripted cases with free exploration. Change the conversation history, the retrieved content and the user role. When something works, run it again a few times. One refusal tells you very little about the next attempt.

  • Log the model version, system prompt, tools and effective permissions for every run.
  • Capture inputs, retrieved content, attempted tool calls and what actually changed.
  • For each finding, note the attacker's access, how to reproduce it, the impact and who owns the fix.
  • Treat the evidence as sensitive. Test logs pick up credentials and customer data.

What should the findings change?

The boundary that failed. Editing the system prompt might change behavior. It does not stop a service account from reading every customer. Apply least privilege to the assistant's identity and check each request against the actual user and record.

AI agent guardrails need to limit execution, so an agent that falls for an attack still cannot send, delete or export much.

  • Enforce document permissions at retrieval, before content reaches the model.
  • Limit tools, arguments and destinations to what the workflow needs.
  • Show approvers the exact target, data and effect of a risky action.
  • Turn every confirmed exploit into a regression test. Rerun it after changes to the model, prompts, retrieval sources or permissions.

How Identra thinks about it

When a red team talks an agent into something harmful, the fix has to live outside the model. On macOS and Windows devices, Identra checks AI agent tool calls against policy and can deny destructive shell commands when policy is set to block. Every agent run is recorded with the user, device, AI client and outcome, so a retest shows whether the attack was actually blocked. In the browser, prompts are checked on the device before they are sent, and sensitive data can be masked or blocked.

Go deeper: AI security, built on identity

Frequently asked questions

Can automated tools replace human red teamers?

No. Automation is good at replaying known attacks after every change. People bring business context and chase failures that cross systems. Use both.

Is the model provider's red team report enough?

It helps. It says nothing about your data sources, permissions, prompts and connected tools. Those create attack paths the provider never saw.

Should you red team in production?

Start in an environment that mirrors production permissions. Production testing needs explicit sign-off, tight limits, monitoring and a way to roll back.

Is a leaked system prompt a serious finding?

Depends what is in it. Generic instructions are low impact. An API key or customer data in the prompt is a real problem.

Who owns the fixes?

The team that owns the boundary that failed. That might be application engineering, identity or the data platform team. A named system owner tracks retests and any accepted risk.

Related terms

Keep exploring · AI security programs and controls