What is AI hallucination risk?

By Identra · Updated

AI hallucination risk is the potential harm from AI output that sounds factual but is invented or unsupported. It becomes a security risk when a person or an automated system trusts that output to write code, change a system or make a consequential decision.

Why do AI models make things up?

A language model predicts likely text. It doesn't look anything up unless a tool or retrieval step hands it a source, and even then it can drift from what the source says.

So you get citations to papers nobody wrote. A Python function that was never in the library. A config flag that sounds right for your version and doesn't exist. The prose is equally fluent either way, and that's the trap. A detailed answer looks researched whether or not anything was checked.

Agents produce a quieter version. When a coding agent reports that the tests pass, that's a claim. You need the test output before it counts.

When does a hallucination become a security problem?

When something downstream acts on it. A wrong sentence in a draft gets caught in review. The same sentence in an incident summary can close a live investigation. A few hypothetical examples:

  • A coding assistant tells a developer to `pip install` a package that doesn't exist. An attacker who registers that name later gets code execution on every machine that follows the advice.
  • A coding agent decides a directory only holds build leftovers and asks to delete it. It held the only copy of a migration script.
  • A support assistant invents a refund exception, and a rep promises it to a customer in writing.

How is hallucination different from prompt injection?

A hallucination needs no attacker. Prompt injection is someone else's text steering the model. Both can end in the same bad action, so find out which one happened before you pick a fix.

Plenty of wrong answers are neither. The model may have accurately summarized a SharePoint page that went stale last year. Or a tool call returned a permission error and the agent carried on with half the data.

  • Hallucination

    Example
    Assistant invents a policy exception
    What to check
    Find the exception in the actual policy
  • Prompt injection

    Example
    A retrieved web page tells the agent to drop its task
    What to check
    Whether outside content is being treated as instructions
  • Stale source

    Example
    Assistant repeats an old retention rule
    What to check
    Whether the source is current and authoritative
  • Execution failure

    Example
    An update fails with access denied
    What to check
    The raw tool result and the real system state

What does it look like during an investigation?

Say an analyst asks an AI assistant about a suspicious Okta sign-in. The assistant finds a legitimate change ticket from the same week and invents a link between that ticket and the account. Its summary calls the activity authorized. If the analyst closes the case on that summary, a real compromise stays open.

The fix is unglamorous. Make the assistant separate what it observed from what it inferred, and attach the record behind each claim. The analyst checks the account, the timestamp and the scope of the approved change. No evidence, no closure.

Agents make this worse because they act on their own conclusions. Excessive agency is what turns a wrong guess into a deleted directory. Valid credentials prove who is acting. They don't prove the action is right.

How do you stop a wrong answer from doing damage?

Work on both ends. Better evidence going in, hard limits on what the output can touch.

For evidence, point the model at current, approved sources and make it cite specific passages. Retrieval helps, but it can't make the model read correctly, and RAG security matters because retrieved pages can carry hostile content of their own.

For output, treat generated commands, code and tool arguments as untrusted input. Skipping that step is what insecure AI output handling describes. Apply least privilege to the agent's credentials and enforce it outside the model, so a mistaken conclusion can't widen its own access.

  • Check unfamiliar packages against the real registry and publisher before installing.
  • Run generated code and risky commands in an isolated workspace first.
  • Validate the target and arguments of every tool call before it executes.
  • Use a dry run when the tool has one, like `terraform plan` before `terraform apply`.
  • Confirm completion from tool output and system state. The agent's summary is not proof.

What should a human reviewer actually see?

The diff. The affected resources, the test output, the source records. Human review that only shows the agent's reassuring explanation checks nothing.

Asking a second model to grade the first can catch mistakes, but it isn't independent proof. Executed tests and authoritative records are.

Log the proposal, the approval, what ran and what came back. When a bad claim gets through, pause the work that depends on it and look at real system state. Then rerun the same task and see whether the revised workflow catches it this time.

How Identra thinks about it

A wrong answer hurts most when an agent acts on it. On macOS and Windows endpoints, Identra checks AI agent tool calls against policy and denies destructive shell commands when policy is set to block. Every agent run is recorded with the user, device, AI client and outcome, so teams can see what the agent actually did instead of trusting its summary.

Go deeper: AI security, built on identity

Frequently asked questions

Is every incorrect AI answer a hallucination?

No. Stale sources, failed tool calls and misread instructions also produce wrong answers. Hallucination usually means content the model invented or couldn't support, presented as fact.

Can retrieval eliminate hallucinations?

No. Retrieved documents give the model context, but it can still misread them or add claims they don't support. Check consequential claims against the source.

Does a confident answer mean the model checked its facts?

No. Tone tells you nothing about verification. Look for sources, test results and observed behavior.

Can a model reliably check its own answer?

Self-review sometimes catches errors. It can also repeat the same bad assumption. Use records or executed tests for anything important.

Is hallucination always a security incident?

No. A made-up claim caught in review is a quality issue. Treat it as a possible security incident when it leads to unauthorized access, unsafe changes or data exposure.

Related terms

Keep exploring · AI threats and attacks