What are computer use agents?

By Identra · Updated

Computer use agents are AI systems that operate a browser or desktop the way a person does, reading the screen, clicking and typing to finish a task. What they can do depends on the apps, files and signed-in accounts in the environment they control.

How do computer use agents work?

You give the agent a goal, say updating a customer record in Salesforce. It takes a screenshot or reads the page structure, picks an action, does it, and looks again. The model decides. A harness around it actually moves the pointer, types and navigates. Anthropic's computer use tool works this way.

That lets an agent use apps built for people, including ones with no usable API. Some agents stay inside a browser. Others drive desktop apps and local files. One task can hop across several apps, so the security boundary is everything the agent can reach. Agentic browser security covers the browser part.

How are computer use agents different from API agents and RPA?

Computer use describes how the agent acts. It says nothing about how smart or autonomous it is. Traditional RPA, like a UiPath bot, replays a predefined workflow. A computer use agent picks its next click from what it sees. Plenty of products mix interface actions with API calls.

Neither path is safe by default. Ask which account authorizes the action, what that account can reach, and which operations need separate approval.

  • Computer use agent

    How it acts
    Reads the interface and picks clicks or keystrokes
    Where the risk sits
    Signed-in sessions, whatever is on screen, unintended actions
  • API agent

    How it acts
    Calls defined operations through tools
    Where the risk sits
    Credential scope, tool permissions, request validation
  • Traditional RPA

    How it acts
    Runs a fixed sequence or rule-based workflow
    Where the risk sits
    Stored credentials, workflow permissions, exception handling

Why are signed-in sessions a security risk?

An agent driving an employee's logged-in browser can read Outlook, download from SharePoint or change settings with no password prompt. The app enforces the account's permissions, and those are usually much wider than the task. Asking for a summary doesn't make the session read-only.

No credential theft is needed. The employee handed it over. MFA protected the sign-in and has nothing to say about what happens after. The app's audit log may show the employee's name with no hint that an agent did the clicking.

AI agent authorization has to weigh the task against the access sitting in the session. A separate Chrome profile isolates cookies. Sign into the same admin account there and the agent has admin rights again.

What can go wrong in an enterprise workflow?

Say a support agent is asked to summarize a customer ticket. The ticket says to open the internal document portal and upload a file to an external verification page. If the agent takes that text as an instruction, it uses real access to leak the file.

That's indirect prompt injection. The ticket's author has no say over the employee's goal, however urgent the text sounds.

A safer setup lets the agent see only the relevant support records and sends its draft to a person. Unrelated domains are blocked. Anything leaving the company needs approval. Those limits have to hold when the model falls for the trick.

Plain mistakes happen too. Wrong recipient. Wrong account. A form submitted twice after a page hung. Depending on the design, screenshots also go to the model provider for processing, so sensitive data can leave before the agent uploads anything.

How do you secure computer use agents?

Start with least privilege and an isolated environment. A system prompt telling the agent to be careful is guidance. Real limits sit outside the model. Anthropic's computer use reference guidance recommends a dedicated virtual machine or container with minimal privileges.

Make human approval specific. The reviewer sees the account, the operation, the destination and the data before anything runs. If any of those change, it needs a fresh approval. The agent never approves its own request.

  • Decide the allowed task, apps, accounts, files and destinations before granting access.
  • Use a dedicated account where you can. Clear out unrelated sessions and admin rights.
  • Limit network access, shared folders, clipboard sharing and downloads to what the task needs.
  • Require approval for external sends, sensitive uploads, payments, deletions and permission changes.
  • Find out where screenshots and task records are processed and how long they're kept.
  • Have a tested stop button, plus a separate way to revoke the sessions and credentials the agent used.

How do you verify the controls work?

Keep an AI audit trail linking the task, the person who started it, the agent, the device, the app account, approvals and what actually happened. Restrict who can read it, and don't keep screenshots you don't need. Check important changes in the destination app itself. The agent's completion message can be wrong.

Test with synthetic data and nasty cases. Hostile ticket text. A fake dialog. A silent account switch. An unexpected recipient. A submission that times out halfway. Denied actions should stay denied, and approval should match the action that actually ran. Hit stop and confirm nothing else happens.

If something goes wrong, stop the run, revoke what it was using, and go through what it already did. New sharing links, changed permissions, submitted forms. Killing the session doesn't undo any of it.

How Identra thinks about it

Identra detects AI agents driving the browser and can block them by policy. Across endpoints and connected identity, SaaS and cloud services, security teams can inventory AI tools and agents, review every endpoint agent run, and have an analyst revoke risky OAuth grants or revoke sign-in sessions with approval.

Go deeper: AI security, built on identity

Frequently asked questions

Does a computer use agent need its own credentials?

No. It can work through a session someone already signed into. A dedicated account with limited permissions makes its access easier to restrict and its actions easier to attribute.

Does MFA prevent an agent from misusing a session?

No. MFA protects the sign-in. Actions taken afterward in an active session need their own authorization and approval controls.

Is a sandbox enough to secure computer use?

No. A VM or container protects the host machine. Accounts signed in inside it keep their full permissions, so accounts and network destinations need limits too.

Can a read-only task still expose sensitive data?

Yes. Depending on the design, screen contents go to a model service for processing. And telling the agent to only read doesn't remove write permissions from the account.

Can computer use agents work without APIs?

Yes. They use visible controls and typing. Many workflows mix interface actions with API calls, so review both paths.

Related terms

Keep exploring · AI apps, agents and usage