What is LLM security?

By Identra · Updated

LLM security is the practice of protecting large language models, the applications built on them and the data they handle from unauthorized access, manipulation, disclosure and abuse. It covers the model lifecycle and the boundaries between prompts, retrieved content, generated output and connected tools.

What does LLM security protect?

An LLM app is a lot more than the model. There's application code, user accounts, data sources and often tools that take action. A model that refuses harmful requests can't make up for an app that lets one user read another user's records.

In scope are training and fine-tuning data, model files, system prompts, retrieved documents, chat history, outputs, API keys and logs. Call a hosted model through the OpenAI or Anthropic API and the provider's data handling becomes part of your review. Self-host Llama on your own GPUs and the serving stack and model files are yours to protect.

What are the main LLM security threats?

Prompt injection is the one everyone's heard of. Instructions get planted in a user message, a retrieved document, a web page or a tool response, and the app follows them. The weights never change.

Most of the rest is ordinary security around the model. A missing authorization check exposes private documents. A leaked API key opens up connected services. Insecure AI output handling is when downstream code renders or runs model output without validating or encoding it. That's how a chatbot ends up delivering cross-site scripting.

  • Data disclosure: sensitive information shows up in answers, shared chats, retrieval results or logs.
  • Tampering: a poisoned dataset or swapped model file changes behavior.
  • Tool abuse: the app takes an unauthorized action through a connected tool.
  • Resource abuse: huge requests or runaway loops burn capacity and money.

How could an LLM application be attacked?

Say a support assistant searches tickets and can email case summaries. An attacker files a ticket telling the assistant to fetch another customer's incident report and mail it to an outside address. The assistant reads that ticket during normal work.

It only lands if the app allows it. Retrieval with too much access lets the model see the report. An unrestricted email tool sends it. Adding a line to the system prompt about ignoring malicious instructions might help a little. It enforces nothing.

The fix is in the plumbing. Scope retrieval to what the requester can read and treat ticket text as untrusted. Check recipients before anything is sent, and for sensitive mail show the exact recipient and content for approval. Then it matters much less whether the model falls for it.

How is LLM security different from AI security and agent security?

LLM security sits inside AI security, which also covers other model types and how a company uses AI in general. Agentic AI security adds delegated authority, memory and actions across systems. The edges overlap a lot.

A locked-down internal chatbot does nothing about an engineer pasting source code into a personal ChatGPT account. And testing a model's answers tells you nothing about an agent's OAuth scopes.

  • LLM application security

    The question
    Can the app expose data or misuse its own output?
    Example control
    Check document permissions before retrieval results reach the model
  • Workforce AI security

    The question
    Which AI services and accounts may touch company data?
    Example control
    Keep sensitive uploads on approved services and work accounts
  • Agent security

    The question
    What can the agent do on someone's behalf?
    Example control
    Limit tool authority and require approval for sensitive actions

How do you secure an LLM application?

Map the users, data sources, model providers, storage and tools. Mark where private data comes in and where output can trigger an action. Every access decision gets an owner.

For retrieval apps, RAG security means source permissions survive indexing, retrieval and caching. Apply least privilege to tool credentials. Enforce policy in application code or in the receiving service, whatever the model thinks.

  • Authenticate users and check authorization before private content reaches the model. Isolate tenants, chat history and caches.
  • Keep credentials out of prompts. Set retention and access rules for conversations and logs.
  • Validate tool arguments against strict schemas. Use parameterized queries and encode output for wherever it's rendered.
  • Tie approval for sensitive actions to the exact action and target, and reject it if either changes.
  • Verify where model files and dependencies came from. Review fine-tuning data.
  • Set limits on requests, execution time, retries and spend. Have a way to switch off a compromised tool.

How do you test that LLM security controls work?

Test the whole app with realistic permissions, retrieval paths and tools. Set up two test tenants and see whether one can pull the other's documents, chats or cached answers. Plant instructions in test documents and tool responses and watch what happens.

AI red teaming should measure harm. A refusal message is not a pass. Try to get unsafe output executed, approval skipped and resources exhausted, using synthetic sensitive data so the test can't cause a real leak.

Rerun after model, prompt, data or permission changes. Failing cases become regression tests. Practice killing a tool and revoking its keys before you need to.

How Identra thinks about it

Identra covers the workforce side of LLM security. Prompts to supported browser AI apps are checked on the device before they're sent, then allowed, masked or blocked by policy. Prompts to Claude Code and Codex are checked on the device before they're sent and blocked when policy is active. Teams can steer ChatGPT, Gemini, Claude, Perplexity and Grok use to company accounts, check AI agent tool calls against policy, and have analysts revoke risky OAuth grants with the result recorded.

Go deeper: AI security, built on identity

Frequently asked questions

Does using a hosted LLM remove the need for application security?

No. The provider secures its model service. User access, the data you send, retrieval, output handling and connected tools are still yours.

Can a system prompt enforce security policy?

No. It can guide behavior, but it isn't an authorization boundary. Enforce data access, tool permissions and approvals outside the model.

Is every hallucination a security incident?

No. Usually it's a reliability problem. It becomes a security problem when a system uses the wrong answer to grant access, run a command or make a consequential decision.

Does retrieval make an LLM application secure?

No. Retrieval brings in content that can carry hostile instructions or records the user shouldn't see. Check permissions and treat retrieved text as untrusted.

Can an LLM application safely execute generated code?

Only in a locked-down sandbox with tight limits on files, network, credentials and resources. Generated code should never run with the app's full privileges.

Related terms

Keep exploring · AI security fundamentals