What security teams miss about coding agents
Your vendor review covered the model. It didn't cover what Claude Code can run, read and install on a developer's laptop.
On this page9 sections
The vendor review is the easy part
Most security reviews of Claude Code or Codex start with the vendor. Data retention, training opt-outs, where prompts get processed. That review is worth doing. It just doesn't touch the laptop.
A coding agent runs as a process under the developer's account. It reads files, runs shell commands and installs dependencies. What it can actually reach depends on the OS, the sandbox, the permission mode the developer picked, and whatever credentials happen to be lying around. That last one is usually where it gets interesting. A .env file with a staging database password. An ~/.aws/credentials file. A GitHub CLI that's already logged in.
The agent doesn't automatically get everything the developer can do. Some clients sandbox file writes. Some ask before every shell command, and some developers turn that off on day two. You have to look at the real configuration on a real machine. OWASP's cheat sheet on secure coding with AI maps out the trust boundaries involved.
So putting "Claude Code" on the approved list doesn't settle much. The same client is fairly harmless in a throwaway container. On a staff engineer's MacBook with a production kubeconfig, it isn't. Coding agent security is about that second case. What work is the agent allowed to do, and what access is sitting there for it to use?
A bug fix that turns into something else
Say a developer on the billing team asks Claude Code to fix a broken invoice CSV export. The repo is checked out locally. There's a .env.local with a staging API key. The AWS CLI is still authenticated from last week's incident. A Jira MCP server lets the agent pull up the ticket.
The agent reads the ticket. Someone left a comment on it suggesting a diagnostic helper from npm to dump the export config. The agent installs it and runs it. A package nobody reviewed has now executed with the developer's full environment, staging key included.
The patch itself is fine. It's a small change to the export formatter, and a teammate approves the pull request on Friday afternoon. Nothing in the diff mentions the npm install or what the helper did with that key.
Run this as a tabletop with your own team. Where should the agent have stopped? At the install, at reading .env.local, or at the AWS CLI? Who would have been asked, and what would they have seen in the prompt? Then try it in a scratch VM with fake credentials and see what actually happens.
One reasonable answer: the agent can edit the export code and run the test suite. A new package or anything that calls AWS needs a person to say yes.
Approving a coding tool means approving what it can reach. Check what that is.
The developer's access is not the task's access
Engineers collect access over time. Write access to a dozen repos, admin on staging, a break-glass production role from the on-call rotation. Fixing a CSV export needs almost none of it.
Split routine development from the operations that hurt when they go wrong. Editing a branch and running tests sit on one side. Publishing to npm, pushing to main and changing a production Terraform variable sit on the other. OWASP calls the gap between what an agent needs and what it holds excessive agency. A perfectly legitimate tool can still be carrying far more authority than the job calls for.
For each workflow, write down where the credential comes from and what it's for. Does the task need a GitHub token, and scoped to which repo? Could the staging test run as a dedicated service account? Can the whole thing run with no production credentials on the machine? Your identity team can review answers like those. They can't review "developers will be careful."
Least privilege only holds if the narrow path is the easy one. Give developers test accounts, small seeded datasets and credentials that expire when the work is done. If the only way to close the ticket is to borrow their personal admin token, they'll borrow it.
Approvals need the same precision. Clicking "allow for this session" at 9am tells you nothing about the upload the agent wants to do at 4pm. A useful prompt names the operation and the target, and says whether it can be undone.
What developers bolt on
An approved-client list is where the inventory starts. Developers keep adding to it. An MCP server pasted into ~/.cursor/mcp.json from somebody's README. A folder of skills dropped into .claude/skills. A plugin someone found in a marketplace thread. Each one changes what the agent can do and what it can touch.
Review each addition against what it's actually for.
- Packages. Ask why it's being added, then check the source, the version and any install scripts. An agent suggesting a plausible package name is not evidence that the package is real, or that it's the one you think it is.
- MCP servers. MCP security starts with plain questions. Who runs the server? What tools does it expose, how does it authenticate, and what can those tools change? A server that reads Jira tickets and one that edits Terraform state need very different permissions.
- Skills and plugins. These carry instructions and sometimes executable code. Decide who can install them and where approved versions come from. An update that asks for new access should go back through review.
- Ownership. Every addition gets a named owner, a stated purpose, a list of what it can reach and a way to remove it. When an unfamiliar MCP server shows up in a session at 2am, that's the first thing a responder will look for.
Text in the repo is not an instruction from you
A coding agent reads far more than the developer's prompt. READMEs, issue comments, code comments, tool output. Anyone who can write to those can try to steer the agent. That's indirect prompt injection, and OWASP's prompt injection cheat sheet covers how it works.
What matters is the point where information turns into authority. A troubleshooting comment can explain a bug. It shouldn't be able to get the agent to read ~/.ssh, paste a token into a gist, or post to some webhook. How official the text sounds is beside the point.
A system prompt telling the agent to ignore suspicious instructions helps a little. Don't lean on it. Limit what the task can reach and where its data can go, and require approval for the sensitive operations. Those limits also cover the more boring failure, where the agent just misreads something harmless.
This is cheap to test. Plant an issue in a test repo asking the agent to print a fake secret. Add a line suggesting it upload a dummy log to a paste site. Watch whether it refuses, asks, or just does it. Look at the approval prompt too. Did it give the developer enough to say no? The result only holds for the setup you tested, so run it again after you change tools or permissions.
A comment in an issue can explain a bug. It can't grant permission.
A checklist that ends in decisions
Get endpoint security, identity and engineering into the same review. "Use AI responsibly" is not something anyone can check when the agent is asking to run terraform apply.
Walk one real laptop workflow through the table. Each row should end as an approved configuration, a fix with an owner, or a written exception. Tie exceptions to a specific workflow and resource. Otherwise a one-off for the payments repo quietly becomes the rule for everyone.
Area
Identity
- What to find out
- Who started the task and which accounts the agent can use.
- Decision to make
- Name an owner. Scope tokens to the repos and environments the work needs.
Area
Files and secrets
- What to find out
- Which directories, .env files and credential files are readable.
- Decision to make
- Give the task its own workspace. Keep unrelated secrets out of reach.
Area
Commands
- What to find out
- Which shell operations the task really needs.
- Decision to make
- Allow build and test. Require approval for destructive or admin commands.
Area
Packages
- What to find out
- What gets added, from which registry, with what install scripts.
- Decision to make
- Review new dependencies before install, in the normal code review.
Area
MCP servers, skills, plugins
- What to find out
- Who owns each one and what it can reach.
- Decision to make
- Approve a purpose, a permission set and an update process.
Area
Outbound destinations
- What to find out
- Where source, logs and diagnostics can be sent.
- Decision to make
- Limit sharing to destinations approved for that data.
Area
Response
- What to find out
- Who can stop a session, revoke access and keep the evidence.
- Decision to make
- Rehearse it once before widening the rollout.
When you need to explain what happened
Something odd happens in an agent session. The responder needs a chain. Who asked for what, on which device, in which session. What the agent tried, what went through, what was blocked and what someone approved. Then the files and remote resources it touched, where you have records for them.
Keep attempted and completed separate. An agent saying "I've rotated the key" is a claim. Go check the destination system. Failures work the same way. A command your policy blocked and a command that died on a typo are different events, and they shouldn't look the same in the record.
Build the AI audit trail around the questions a responder will ask. Who authorized it, what access got used, what changed, what has to be rolled back. You don't need full repo snapshots or secret values in the log to answer those. Storing them creates a new sensitive dataset with its own retention problem.
Rehearse the stop. Killing the agent process doesn't revoke the GitHub token it used, or an OAuth grant it picked up through an MCP server. Decide ahead of time who handles revocation in GitHub, in AWS and in Okta, because those may be three different teams.
For rollout, start narrow. Local test repair on synthetic data is a decent first workflow. Widen it when there's a business reason and someone can explain how sensitive actions get approved and how a bad session gets investigated.
From Identra
How Identra helps
Identra helps teams secure coding agents on macOS and Windows. It finds coding agents such as Claude Code and Codex on each machine, along with the MCP servers, skills, plugins, packages and agent tasks around them. Prompts to Claude Code and Codex are checked on the device before they're sent, and blocked when policy is active. Agent tool calls are checked against policy. Destructive shell commands are denied when policy is set to block. On macOS, access to protected files can be denied. Every agent run is recorded with the user, device, AI client and whether it was allowed or blocked. Where the person is known, their browser, endpoint and provider activity sits on one timeline.
Frequently asked questions
Does a coding agent get all of a developer's permissions?
Not automatically. It depends on the sandbox, the permission mode, and which credentials are on the machine. Running under the developer's account is a reason to check, not proof that everything is reachable. Look at the actual setup.
Isn't code review enough?
No. The pull request shows the final diff. It doesn't show the package the agent installed along the way, the .env file it read, or the remote call it made while troubleshooting. Review those alongside the patch.
Should we block all MCP servers?
Blocking everything mostly pushes developers to work around you. Review each server for who runs it, what tools it exposes, how it authenticates and what it can change. Keep the ones with a clear purpose and an owner. Remove the ones nobody can justify.
Where do we start?
Pick one routine task, like fixing a failing test, and walk through it on a representative developer laptop. List the files, credentials, commands and connected tools the agent can reach. Cut what the task doesn't need and decide which actions need a separate yes.
Related terms
Solutions
See it in your environment
AI security built on identity. One timeline.
A walkthrough with a security engineer.
