What is AI agent skills and plugins security?

By Identra · Updated

AI agent skills and plugins security is the practice of reviewing and controlling the instructions, code, and connections that extend an AI agent's abilities. It limits what those additions can access and do, including how they use permissions the agent already holds.

What are AI agent skills, plugins and extensions?

A skill is a folder of instructions, built around a SKILL.md file and sometimes bundled with scripts, templates or reference docs. The agent loads it when a task calls for it. The format is set out in the Agent Skills specification. In Claude Code, skills can live in your home directory under ~/.claude/skills or inside a repository's .claude/skills folder.

Plugin and extension mean whatever the host application says they mean. They might bundle code, add tools or connect an outside service. A plugin that wires in an MCP server can hand the agent a new set of tools in one install.

A text-only skill gets no new permissions. It can still tell the agent how to use the access it already has, and that's enough to matter.

  • Skill

    Typical purpose
    Task instructions plus supporting files
    Main review question
    What does it ask the agent to do, and what scripts can it run?
  • Plugin or extension

    Typical purpose
    Code, integrations or tools
    Main review question
    What executes, and which data or accounts can it reach?
  • MCP server

    Typical purpose
    Tools, resources or prompts exposed over MCP
    Main review question
    Which capabilities are exposed, and how is access authorized?

Why can a new skill change an agent's risk?

It can change behavior without changing permissions. A reporting skill tells the agent to gather files it can already read. A publishing plugin lets it post them somewhere. Install both and you've built an exfiltration path out of two reasonable tools.

Work out whose authority each action runs under. The developer's own session? A dedicated agent account? A token shared across scheduled jobs? Apply least privilege at that level, and look at OAuth app risk when a connector holds delegated access to Google Drive or Slack.

A skill that says never read .env files is making a request of the model. A file permission or sandbox rule is a boundary. You want both.

How can skills and plugins become attack paths?

A malicious skill can ask for credentials, a shell command or an upload and attach a plausible reason, like collecting diagnostics for the support team. A legitimate one can pull in untrusted content, an issue body or a web page, that carries prompt injection.

Bundled scripts bring the usual software problems, from vulnerable dependencies to hijacked updates. Text changes count too. An edit to SKILL.md can swap a destination URL without touching a single line of code.

Publisher identity, source integrity, dependencies and update control all fall under AI supply chain security. Review tells you what you installed. Runtime limits cover whatever the review missed.

What does a risky workflow look like?

Say an engineering team installs a release-notes skill. Their coding agent can read the repository and run shell commands. The skill ships a helper script that gathers diagnostic files and posts them to an outside support endpoint.

Release notes need commit messages and issue titles. If the script also grabs .env and local config, secrets leave through access the agent already had. No new permission prompt appears, as long as existing settings allow the reads and the upload.

A safer setup feeds the skill approved release metadata only, keeps credentials out of the working directory and restricts outbound destinations. If an upload really has to happen, the approval screen shows the exact files and URL. Then someone tests that the agent can't read the forbidden files at all.

How should security teams review and restrict additions?

Approve a specific version, for a specific task, owner and environment. Approving Cursor or Claude Code doesn't approve everything they can load.

Use agent sandboxing with explicit limits on files, credentials, commands and network destinations. A sandbox with open outbound access won't stop an upload.

  • Record the publisher, source, version or content hash, business owner and intended use.
  • Read the instructions. Inspect bundled scripts, dependencies, install steps and any remote URLs.
  • Map each capability to the account, credential, files and service permissions it can use.
  • Split read access from write, delete, publish and admin actions wherever the service allows it.
  • Test with fake secrets and hostile tool output. Confirm the forbidden reads, uploads and commands fail.
  • For consequential actions, show the real target, data and change before anything runs.

How do you manage updates and remove access safely?

Keep an inventory of enabled additions with owners, versions, permissions and connected services. Pin versions where the host supports it so an approved package can't change underneath you. Reassess when instructions, code, publisher, destinations or grants change.

Log the acting identity, action, target, approval and outcome in an AI audit trail, with secrets kept out. Check that responders can actually disable an addition and stop work already in progress.

Uninstalling removes the local files. It doesn't revoke an OAuth grant the plugin created, a token it stored or a scheduled job it set up. Check each connected service, and rotate secrets if exposure is suspected.

How Identra thinks about it

Identra inventories skills, plugins, MCP servers and IDE extensions on endpoints, alongside browser extensions and connected AI apps with access to your tenant. Teams can enforce policy on AI agent tool calls and disable risky browser extensions. Analysts can revoke risky OAuth grants.

Go deeper: Identra on the endpoint

Frequently asked questions

Can a skill be dangerous if it contains only text?

Yes. The instructions shape how the agent uses tools and permissions it already has. Read them even when there's no code.

Does marketplace approval make a plugin safe?

It's useful evidence. It doesn't tell you whether the plugin's access fits your task. Check the source, permissions, data handling and how updates arrive.

How are skills different from MCP servers?

A skill is mostly instructions and supporting files. An MCP server exposes tools over a protocol. A skill can tell the agent to call an MCP tool, so review what they can do together.

Can a skill's permission declaration enforce least privilege?

Only if the host enforces it. Treat a declaration as a claim and verify it against host settings, operating system restrictions and service authorization.

Does removing a plugin revoke its credentials?

Not necessarily. OAuth grants, stored tokens and scheduled jobs can survive removal. Confirm revocation in each connected service.

Related terms

Keep exploring · AI apps, agents and usage