What is AI incident response?
By Identra · Updated
AI incident response is the process of investigating, containing and recovering from security incidents that involve AI applications, agents, models or the systems they connect to. It establishes what the AI received, what it actually did and whose access it used, so responders can stop the harm and restore trusted operation.
How is AI incident response different from normal incident response?
The playbook is the one you already run. Triage, contain, preserve, eradicate, recover. What changes is where the evidence lives. You now need the prompt, the documents the model pulled in, every tool call it made and anything it wrote to memory. Then you have to tie all of that to effects in real systems.
The bigger shift is that nobody has to steal a password. Say a coding agent reads a web page carrying a prompt injection payload, then uses its own legitimate GitHub token to push a change. Every log line shows a valid, authenticated actor. Authentication tells you who acted. It does not tell you the business wanted it done.
Scope note. This entry is about incidents that involve AI. Using an LLM to summarize alerts for your SOC is a different topic.
What should responders establish first?
Get the basics written down before anyone starts revoking things. Which workflow, who owns it, which person or agent kicked it off, which account it ran as. Was this ChatGPT in a browser tab, Claude Code on a laptop or a Copilot Studio agent running in the tenant? Set the time window.
Then separate what happened from what could have happened. An agent with read access to a repository did not necessarily read every file in it. For suspected AI data leakage the questions are narrow. Was the data actually sent, to which service, under which account, and who can see it now? A paste that was blocked before sending and a finished upload to a personal account are different incidents.
Customer export uploaded to a personal ChatGPT account
- What to establish
- File contents, the receiving account, whether the chat was shared
- First move
- Stop further uploads and start the provider's deletion process
Coding agent ran an unexpected shell command
- What to establish
- The input that triggered it, the exact command, the identity it ran as, what changed
- First move
- Pause the agent and rotate any secrets it could read
Unfamiliar OAuth app holding Mail.Read across mailboxes
- What to establish
- Who consented, which scopes, whether the app has used them
- First move
- Revoke the grant, then review mailbox activity
What does an AI agent incident look like?
Say a support agent reads incoming tickets and can look up customer records and send email. An attacker files a ticket telling it to pull records for unrelated customers and mail them to an outside address. The agent treats the ticket text as instructions. It calls its tools.
Responders pause the agent first. Then they save the ticket, the run history, the tool requests and the destination address. The CRM's own logs show whether records were retrieved, and the mail system's logs show whether anything left. The agent's own message saying it finished the task proves neither.
If records went out, the data owners get pulled in to work out which customers are affected. Before the agent comes back online someone clears the malicious ticket from its queue, checks its memory for planted instructions and cuts its send permission down to internal addresses.
How do you contain an AI incident and revoke access?
Stop the harm while you collect what you can. Do not leave a dangerous agent running so the timeline is complete. Containment has to reach every system the AI can touch. Blocking chatgpt.com in the browser does nothing about a connected app that already holds a token to your Google Drive.
Handle OAuth app risk as its own track. Revoking a grant, resetting a password and ending a session do different things at different providers. Confirm what each action actually ended, then watch the resource for more activity.
- Pause agent runs, schedules and delegated tasks. Look in downstream queues for work it already submitted.
- Disable the browser extension or remove the MCP server entry, for example from ~/.cursor/mcp.json. Isolate the laptop if the evidence points past the AI tool.
- Revoke the grants and sessions involved and rotate exposed secrets. Write down the target and result of each action.
- Restrict shared conversations. Ask the provider what retention and deletion options exist for that account type.
- Check destination systems for changes that already landed. Removing the agent does not undo them.
What evidence should an AI incident response team preserve?
Build an AI audit trail that runs from input to identity to action to outcome. Conversation and run IDs, prompts, retrieved content, tool arguments and results, permission changes, provider audit logs. Grab the model and app configuration too if you can still get it.
Timestamps need time zones. Mark which tool calls were attempted and which executed.
Keep originals locked down with a note on how they were collected. Tickets get references or redacted excerpts. Nobody should be pasting API keys and customer rows into the incident Slack channel.
Write down the gaps. Missing logs mean you do not know what happened, and the timeline should say so.
How do you recover and prepare for the next AI incident?
Restore from a known good state. Before the workflow resumes, review changed files, agent memory, scheduled tasks and connected tools. Issue new credentials with least privilege and replay the original attack path in an isolated environment. One clean run is weak evidence that the hole is closed.
Agree the restart bar with the workflow owner. The harmful activity has stopped, access is under control, unauthorized changes are dealt with and monitoring would catch a repeat. Bring it back gradually with someone watching.
Most of the work happens before the incident. Keep a current list of workflow owners and provider contacts. Practice the pause and revoke steps. Rehearse a scenario that crosses browser, endpoint and cloud, such as a poisoned document that leads a coding agent to commit a live key. The NIST Generative AI Profile recommends defined response ownership, rehearsals and improvement from retrospectives for third-party generative AI.
How Identra thinks about it
Identra puts browser, endpoint and connected-provider activity on one timeline tied to the person behind it, where identity resolves. Every AI agent run on an endpoint is recorded with the user, device, AI client and whether it was allowed or blocked. Analysts can revoke risky OAuth grants and, with approval, sign-in sessions, and every response records its result.
Go deeper: AI security, built on identity
Frequently asked questions
What should you do when someone pastes sensitive data into AI?
Find out whether it was sent, which account received it and what it contained. Stop further exposure, preserve evidence and rotate any secrets. Deleting the conversation may not remove every retained copy, so check the provider's retention and deletion options.
Can an AI agent be compromised without stolen credentials?
Yes. Malicious content can steer an agent into misusing permissions it legitimately holds. Investigate what it read and which tools it called, alongside the authentication records.
Is an incorrect AI answer a security incident?
Not by itself. It becomes one if it exposed protected information, triggered an unauthorized action or caused another security impact. Other failures belong in a quality or safety process.
Does a password reset stop an AI application's access?
Do not count on it. Connected apps can hold OAuth grants, refresh tokens or API keys with their own lifecycles. Revoke the specific access and confirm the result in the provider's controls and the resource's activity.
Who should own AI incident response?
The security incident response team coordinates, working with the AI workflow owner and the relevant administrators. Data owners and privacy or legal join when sensitive information or reporting decisions are involved.
