What is AI agent hijacking?

By Identra · Updated

AI agent hijacking is an attack that redirects an AI agent from its authorized task toward an attacker's goal. The attacker manipulates the agent's instructions or compromises its tools or configuration, and the agent then misuses access it legitimately holds.

How does AI agent hijacking work?

An agent reads content, picks actions and calls tools with access someone gave it. Hijacking means an attacker is now steering those choices. Downstream services see a valid token and a normal API call. Nothing looks off to them.

The usual route is indirect prompt injection. Text in a shared Google Doc, a web page or an MCP tool response gets read as orders. Other routes go through the agent's setup instead: a tampered tool description, an edited agent config file, or stolen credentials that let someone rewrite a scheduled job.

The OWASP AI Agent Security Cheat Sheet calls the objective-level version goal hijacking. An attacker doesn't need lasting control. Redirecting one action can be enough.

How is agent hijacking different from prompt injection?

Prompt injection is a technique. Hijacking is the outcome. One often causes the other, and neither needs a stolen password.

  • Agent hijacking

    What it describes
    An agent redirected toward an attacker's goal
    Example
    A support agent exports records instead of summarizing a ticket.
  • Prompt injection

    What it describes
    Malicious instructions that sway model behavior
    Example
    A retrieved page tells the agent to drop its assigned task.
  • Credential theft

    What it describes
    Someone takes a secret or token
    Example
    An attacker copies the agent's API token and calls the service directly.
  • Excessive agency

    What it describes
    More tools, permissions or autonomy than the task needs
    Example
    A summarization agent that can also delete records and email outsiders.

What does agent hijacking look like in an enterprise?

Say a support agent is asked to summarize a new Zendesk ticket. The ticket claims a compliance audit needs records for several other customers sent to an outside address.

If the agent goes along with it and uses its CRM and email tools, the attacker gets an export they could never run themselves. The ticket submitter has no database access. The agent supplies it.

Whether that works depends on controls outside the model. A CRM query scoped to the ticket's customer limits what comes back. An email tool that only sends to approved domains blocks delivery. A reviewer who sees the actual records and recipient can say no.

Excessive agency widens the damage. A browser agent acts through whatever sessions are signed in. A coding agent runs shell commands. Read-only access still leaks if the agent has any way to send output somewhere.

What are the warning signs of a hijacked agent?

Compare what was asked with what the agent read, changed and sent. A valid login and a successful API response say nothing about intent.

Bugs and vague requests produce the same symptoms. Suspicious input right before the deviation is the stronger tell.

  • Reads of records or files unrelated to the task.
  • A new external recipient or network destination in a tool call.
  • Requests for secrets, wider scopes or an exception to approval.
  • Tool definitions, stored instructions or scheduled jobs changed without the owner knowing.
  • Actions under the agent's identity with no matching task or approval.

How can security teams prevent or limit agent hijacking?

Assume the agent will get fooled at some point. Then make being fooled cheap. Apply least privilege to its tools and credentials and enforce it in the systems that execute the actions. A system prompt is not an authorization layer.

Put human approval on consequential actions and show the exact target, recipient and data. If the destination or payload changes after approval, it goes back for review.

  • Give each agent an owner and one job. Strip the tools outside it.
  • Validate tool arguments against allowed resources and destinations.
  • Keep long-lived credentials out of model context and issue short-lived, scoped ones.
  • Review new tools, plugins and persistent instructions before sensitive workflows use them.
  • Feed the agent malicious documents in a test environment and confirm enforcement holds when it tries to comply.

What should teams do after suspected agent hijacking?

Stop the run and block further tool calls. Preserve the inputs, tool versions, config and action log. Revoke exposed tokens and sessions. Killing the process doesn't invalidate a token that was copied somewhere else.

Then trace what already finished. Look for delegated jobs still running, and for instructions sitting in memory or a schedule that would start the behavior again. Work out which data went where. Run it through your AI incident response process and write down plainly what the logs can't confirm.

How Identra thinks about it

Identra shows security teams which AI apps and agents are in use across the browser, the endpoint and connected identity, SaaS and cloud providers. Teams can block AI agents driving the browser and destructive endpoint shell commands by policy, review recorded agent runs and revoke risky OAuth grants, with the result of each response recorded.

Go deeper: AI security, built on identity

Frequently asked questions

Does agent hijacking require stolen credentials?

No. An attacker can talk an agent into misusing access it already has. Stolen credentials are a separate route when they allow changes to the agent's tasks or configuration.

Is using a stolen agent token always agent hijacking?

No. Calling a service directly with the token is credential abuse. It's hijacking when the attacker redirects or controls the agent itself.

Can a read-only agent still cause harm?

Yes. Read-only protects the source data from changes. The agent can still disclose it in a reply or through another tool.

Can a system prompt prevent agent hijacking?

It can spell out the task and what to distrust. It can't guarantee the model listens. Authorization checks have to limit what happens when it doesn't.

Can hijacking persist after an agent restarts?

Yes, if the malicious instructions sit in persistent memory, config or a recurring task input. Restarting without checking that state brings the behavior back.

Related terms

Keep exploring · AI threats and attacks