What is indirect prompt injection?

By Identra · Updated

Indirect prompt injection is an attack that embeds malicious instructions in content an AI system reads, such as emails, documents, web pages or tool results. It tries to make the AI treat that content as authority and redirect its responses or actions away from the user's authorized task.

How does indirect prompt injection work?

The attacker never talks to the AI. They leave text where it will be read during normal work. A web page a research assistant fetches. A GitHub issue a coding agent opens. An email Microsoft 365 Copilot is asked to summarize. A result an MCP tool returns.

The useful content and the hostile instructions arrive together. The attack works when the model treats them as orders. A document might pretend to be internal policy, ask for something unrelated or tell the assistant to keep a step from the user. It's a form of prompt injection.

Hidden text helps attackers, but they don't need it. Plain visible sentences can do the same job. OWASP's prompt injection guidance covers the indirect case as instructions arriving through external sources the model processes.

How is indirect injection different from direct injection?

Delivery path. Direct injection comes in through the chat box. Indirect injection rides in on content the AI consumes while doing its job, and the user who asked for a summary usually has no idea.

A trusted source can still return attacker text. Your SharePoint search is authenticated. The file it returns may have been edited by a contractor or copied from the internet. That's why RAG security treats retrieved passages as untrusted.

  • Delivery

    Direct injection
    A message typed into the AI
    Indirect injection
    Content the AI retrieves or opens during a task
  • Attacker controls

    Direct injection
    Their own message
    Indirect injection
    Some content the AI happens to read
  • Example

    Direct injection
    A chat message tries to override the app's rules
    Indirect injection
    A retrieved document tells the assistant to share private files

What damage can it do?

With no tools at all, an assistant can still be steered into a slanted summary, a dropped warning or a recommendation the attacker picked. Nobody steals anything. A bad business decision gets made.

Tools raise the stakes. Now the agent can send mail, edit records or run commands. When an attacker gets an authorized agent to misuse its own access, that's the confused deputy problem. The attacker gets the action without ever holding the credentials.

Authentication tells you who is calling. It says nothing about whether this call matches what the user asked for.

What does an attack look like at work?

Say a procurement assistant is asked to summarize a vendor email. The email says an internal compliance review requires finding the confidential pricing sheet and forwarding it to an outside address. That text is part of what's being summarized. The employee never asked for it.

If the assistant can read the pricing sheet and send email, a reading task turns into a leak. The fix sits outside the model. The send should fail because the summary task never authorized sharing a different document.

Same pattern with code. A comment in a repo asks the coding agent to upload .env files for troubleshooting. Look at the source, destination and action together. Words like compliance or maintenance don't authorize anything.

How do you defend against it?

Label retrieved content and keep its source attached. Tell the model how to treat instructions inside it. Then assume that will sometimes fail, and enforce access and execution rules outside the model.

Scope each workflow with least privilege. A summary needs read access. Sending or changing anything should need separate authorization. OWASP's prevention cheat sheet recommends separating untrusted content, restricting tool permissions and requiring approval for consequential actions.

  • Limit mailboxes, repos and document libraries to what the task needs.
  • Validate tool arguments against allowed targets and destinations before they run.
  • Restrict outbound network access and external sharing where you can.
  • Require approval for external sends, destructive commands and permission changes. Show the approver the real recipient and data, and bind the approval to that exact action.
  • Log tool requests and whether they were allowed, blocked or completed.

How do you test your defenses?

Put the hostile text where the assistant actually reads it. A document, an email, a tool result. Typing it into the chat tests a different door.

Decide the expected result first. In the procurement case, the summary gets written and nothing else happens. Check the tool calls as well as the answer. A polite refusal in the chat doesn't mean a send didn't fire earlier.

Try false policy claims and instructions pasted into internal docs. See whether your AI agent guardrails still hold when the model goes along with the request. Rerun after any change to tools, access, sources or the model.

How Identra thinks about it

Identra helps limit what a redirected agent can do. On endpoints, AI agent tool calls are checked against policy and destructive shell commands are denied when policy is set to block. In the browser, AI agents driving the browser are detected and can be blocked by policy. Across connected identity, SaaS and cloud services, analysts can revoke risky OAuth grants so a manipulated agent holds less access.

Go deeper: AI security, built on identity

Frequently asked questions

Can internal documents carry indirect prompt injection?

Yes. A malicious insider can plant instructions, or someone can paste them in from outside. Living on SharePoint doesn't make text trustworthy.

Is every instruction in a document a prompt injection?

No. Docs are full of procedures and example commands. The problem is text that tries to push the AI beyond the task it was given.

Does read-only access prevent damage?

It protects the source systems. The assistant can still give misleading answers or leak what it reads through any output channel it has.

Can a system prompt prevent indirect prompt injection?

It helps the model handle untrusted text. It can't be your only control. Permissions and action checks have to hold when the model gets it wrong.

What should a team do after a suspected indirect prompt injection?

Pause the workflow and keep the source content, tool requests and results. Find out whether data left or anything changed, then contain the access and retest before restarting.

Related terms

Keep exploring · AI threats and attacks