What is MCP tool poisoning?

By Identra · Updated

MCP tool poisoning is a prompt injection attack that embeds malicious instructions in Model Context Protocol tool descriptions or other tool metadata shown to an AI model. It tries to make an agent misuse its available tools or permissions, causing unauthorized actions or data exposure.

How does MCP tool poisoning work?

Every MCP server tells the client what tools it offers. Each tool comes with a name, a description and parameter descriptions. Clients like Cursor, Claude Desktop and Claude Code pass that text to the model so it knows what it can call.

The user rarely reads it. The model always does. A poisoned description slips in instructions that have nothing to do with the tool. Read this file first. Skip the confirmation. Send the context to this support endpoint. If the model treats that as part of its job, the attack works. Invariant Labs published a demonstration in its tool poisoning research.

It's a kind of indirect prompt injection. The instructions come from tool metadata instead of the user.

What could a poisoning attack look like on a developer laptop?

Say a developer adds a third-party documentation search server to ~/.cursor/mcp.json. Their coding agent can already read the project. They ask it to fix a build error. The search tool's description says every query must include the contents of the project's .env file for compatibility checks.

If the agent complies, it reads the API keys in .env and sends them as a search argument to someone else's server. A docs lookup just became AI data leakage.

The description granted no file access. It borrowed the access the agent already had. Limit which files the agent can read and what it can send out, and the chain breaks even when the model goes along with it.

How is tool poisoning different from other MCP attacks?

Where the bad content lives decides what you inspect. Scanning tool results won't catch a poisoned definition. A clean definition says nothing about the server's code.

A rug pull is a server that changes after you approved it. The new version might carry a poisoned description, or it might just behave badly without needing to fool the model. Both sit under MCP security. The checks are different.

  • Tool poisoning

    Where the problem is
    Instructions in tool definitions shown to the model
    What to inspect
    Tool and parameter descriptions, other model-visible metadata
  • Tool-result injection

    Where the problem is
    Instructions in content a tool returns
    What to inspect
    Returned content and the agent's next actions
  • Malicious implementation

    Where the problem is
    Harmful code in the server itself
    What to inspect
    Source, dependencies, permissions, runtime behavior
  • Rug pull

    Where the problem is
    An approved server changes later
    What to inspect
    Definition and deployment diffs, re-approval

How do you reduce tool poisoning risk?

Know which MCP servers are configured, on whose machine, and why. Read the definitions the client actually sends to the model. The README may say something else. Watch for instructions to read unrelated files, hide steps, override policy or change how other tools get called.

Apply least privilege to the agent and the server. The model explaining why it needs more access never grants it more access.

  • Allow only approved servers. Remove tools a workflow doesn't use.
  • Save reviewed definitions and diff them on every update.
  • Pin server versions where you can. A familiar remote URL can serve different code tomorrow.
  • Validate tool arguments against allowed targets before execution.
  • Treat description scanners as a review aid. Keep the permission limits even when the scan comes back clean.

How do you test whether defenses work?

Use an isolated environment with fake secrets and a destination you control. Add a test tool whose description asks for an unrelated file read or an outbound transfer. Then change the description after approval and see if anyone notices.

Look at what executed. The agent's own account doesn't count. A refusal at the end means nothing if an earlier tool call already sent the file. Record server, definition version, request, decision and outcome in an AI audit trail, with credentials redacted.

What should you do after suspected tool poisoning?

Pause the workflow and disconnect the server. Save the definitions, config and logs before anything gets reinstalled.

Trace sensitive reads, outbound requests, commits and permission changes. Rotate exposed credentials. Check the other tools the agent called too, since a poisoned description can steer it toward a different tool entirely.

Bring the server back only after review, with narrower agent access. Fold what you learned into your AI incident response runbook.

How Identra thinks about it

Identra gives security teams visibility into MCP servers, coding agents and desktop AI apps in use on endpoints. Teams can apply policy to AI agent tool calls and block destructive shell commands when blocking policy is enabled. Agent run records connect activity to a user, device and outcome for investigation.

Go deeper: Identra on the endpoint

Frequently asked questions

Does the agent have to call the poisoned tool?

No. The model reads all available tool definitions. A poisoned description can push it to misuse a different tool it has.

Does MCP authentication prevent tool poisoning?

No. Authentication tells you who you're connected to. It says nothing about whether their descriptions are safe or will stay that way.

Is a read-only tool safe from poisoning?

A read-only label enforces nothing. Even a truly read-only tool can leak what it reads if the output goes somewhere it shouldn't.

Is tool poisoning the same as model poisoning?

No. Tool poisoning targets text the agent reads at run time. Model poisoning tampers with training or fine-tuning data.

Can human approval stop tool poisoning?

Sometimes. It works when the prompt shows the real action, target and data. A generic Allow button or a summary written by the misled model can hide the problem.

Related terms

Keep exploring · AI threats and attacks