What is multi-agent system security?
By Identra · Updated
Multi-agent system security protects workflows where AI agents pass information, delegate tasks and act through tools. It controls the authority each handoff carries and keeps untrusted content from turning into an instruction as it moves from agent to agent.
Why do multi-agent workflows need their own security controls?
An orchestrator splits a request across a research agent, a planner and an executor. Each one has its own tools and credentials. The weak spot is usually the handoff. A read-only research agent can still cause damage if the executor acts on its recommendations without checking anything.
A message can come from a legitimate agent and still carry hostile text copied out of a PDF. A valid task can drift too, as agents widen its scope or keep retrying. Controls have to travel with the task.
Authority
- One agent
- Limit its tools and resources
- Several agents working together
- Limit what each handoff can pass along
Untrusted content
- One agent
- Validate what it retrieves
- Several agents working together
- Keep the source attached through summaries and forwards
Containment
- One agent
- Stop it and revoke its access
- Several agents working together
- Also cancel dependent tasks and queued actions
Accountability
- One agent
- Log its requests and tool calls
- Several agents working together
- Link those logs across the whole delegation chain
How should agents establish trust at each handoff?
Give every agent that acts on its own a verifiable identity. A name in a message, or a claim to be the orchestrator, proves nothing. AI agent identity gives the receiving service a principal it can check.
Authentication tells you who sent the request. Whether the request is allowed is a second check, and it belongs in the service doing the work. Permission to call an agent shouldn't unlock everything that agent can do.
Keep source content separate from task instructions as it moves. Summaries should carry their source references. Labels help. On their own they won't stop a model from obeying hostile text.
How should delegated permissions work?
AI agent delegation should name the task, the resources, the allowed actions and when the authority expires. A sub-agent gets what its piece needs. Pass the orchestrator's broad token down the chain and every sub-agent becomes as powerful as the top.
Enforce least privilege at the receiving service. An action has to fit the delegated scope and the receiver's own permissions. Skip that and a low-privilege agent can talk a high-privilege one into acting for it. That's the confused deputy problem with an LLM in the middle.
- Decide whether an agent may delegate again and what it may pass on.
- Use short-lived credentials with narrow resource and recipient limits.
- Recheck authorization when work executes, including after it sat in a queue.
What does a multi-agent attack look like in an enterprise?
Say a procurement workflow has three agents. One reads supplier proposals. One recommends a vendor. One updates records in SAP Ariba and sends email.
A malicious proposal tells the reader agent to attach the internal pricing sheet to the supplier follow-up. The reader passes it on as a recommendation. The planner turns it into a task. The executor sends the file with its own valid access. That's prompt injection picking up authority at every hop.
The fix is unglamorous. Keep the proposal's origin attached. Treat a recommendation as input to check. Have the email service refuse confidential attachments to outside recipients unless someone approved that exact send, even when the request comes from the expected planner.
How do you contain failures and test the controls?
Honest mistakes travel the same paths as attacks. A wrong result lands in shared memory, kicks off dependent tasks or gets retried until a payment goes out twice. Treat shared state as a security boundary.
AI agent sandboxing handles local execution. Remote API permissions and agent-to-agent calls need limits of their own.
- Inventory agents, owners, tools, credentials and allowed handoffs.
- Cap delegation depth, retries, runtime and spend.
- Make retries safe. Check whether a payment or account creation already went through before trying again.
- Build a stop button that cancels dependent work and pulls access.
- Test forged sender claims, stale approvals, poisoned shared memory and cancellation while work is queued.
What should a multi-agent audit trail record?
An AI audit trail links the original request through each delegation to the final action. Record the trigger, the parent and child agents, the delegated scope, authorization decisions, approvals and results. A shared trace ID ties records together. It doesn't prove anything was authorized.
Scheduled workflows often have no human requester. Record the trigger and the owner instead of inventing one.
During an investigation, separate what completed from what is still queued. Revoking a credential doesn't unsend an email.
How Identra thinks about it
Identra discovers AI agents across Microsoft 365 Copilot and Foundry, Google Workspace and Anthropic along with their owners, and ties browser, endpoint and provider activity to one person where identity resolves. Endpoint agent runs are recorded with the user, device, AI client and allowed or blocked outcome, which gives investigators a place to start across agent workflows.
Go deeper: AI security, built on identity
Frequently asked questions
Is multi-agent security different from securing individual agents?
It adds the relationships. Agents that are each locked down can still form an unsafe workflow if handoffs widen authority or lose track of where content came from.
Does encrypted agent-to-agent communication make a workflow safe?
Encryption protects the message in transit. The receiver still has to authenticate the sender, authorize the operation and validate the content.
Should sub-agents inherit the orchestrator's permissions?
No. Give each sub-agent what its assignment needs. Broad inheritance means one misled sub-agent can do whatever the orchestrator can.
Can a supervisor agent serve as the security boundary?
It can review plans, and it can also be fooled. Put access limits and approvals in the services that execute actions.
Does human approval secure every downstream action?
Only the specific action that was approved. Signing off on a goal at the start doesn't cover everything the agents propose later.
