What is model extraction?
By Identra · Updated
Model extraction is copying a model's behavior without authorization by collecting its responses and training a substitute on them. Model theft is the wider term and also covers directly stealing the weights, the learned parameters a model uses to make predictions.
How does model extraction work?
The attacker sends inputs, saves the outputs and trains a copy on those pairs. Access might come through a paid API, a stolen customer account or an app that shows predictions. Often they only want one narrow skill, like the way a model sorts documents into categories.
What comes back matters. Labels, generated text and confidence scores each give away different amounts. Hiding the scores doesn't close the door. The USENIX research on model extraction showed attacks on some prediction APIs that work from labels alone.
Weight theft takes a different route. Someone copies the model files out of an S3 bucket, a model registry, a backup or a laptop. Checkpoints and private LoRA adapters count too. No queries needed.
Why does model extraction matter to enterprises?
A model can carry a lot of investment in data, expert labeling, training and evaluation. A decent replica lets someone else offer something similar without paying for any of it. How much that hurts depends on what was copied, how good the copy is and whether the capability is commercially sensitive.
First work out what you actually own and control. For a hosted model, who controls inference access, and can private artifacts be exported? For self-hosted models, put registries, deployment packages, adapters and backups into AI supply chain security reviews.
A behavioral copy doesn't come with the training data. Or with the retrieval layer and tools wrapped around the model. Lost capability and data exposure are separate questions.
How is model extraction different from data or prompt leakage?
They overlap, but the evidence and the containment differ. AI data leakage is about exposed information. System prompt leakage is about hidden instructions. Neither one proves anybody walked off with the model.
Model extraction
- What the attacker gets
- A substitute that mimics some of the model's behavior
- Evidence to look for
- Systematic response collection, plus signs of a replica
Weight theft
- What the attacker gets
- The actual parameters or private adapters
- Evidence to look for
- Unauthorized reads, downloads or exports of model files
Data leakage
- What the attacker gets
- Confidential records or documents
- Evidence to look for
- Sensitive content in responses or outbound transfers
System prompt leakage
- What the attacker gets
- Hidden application instructions
- Evidence to look for
- Responses that reveal protected instructions
What does a model extraction attack look like in practice?
Say a vendor sells a support-ticket classifier through an API. One customer account sends carefully varied tickets and stores every category that comes back. Then it trains a competing classifier on those labels. That's extraction, and nobody ever touched the vendor's model files.
The vendor might notice steady collection across related accounts, or requests that keep probing the edges between categories. Legitimate evaluation can look exactly the same. Check account ownership, the terms of use, request history and any evidence of a substitute before calling it theft.
Different case. A deployment credential can download the classifier's checkpoint from storage. Someone steals the credential and exports the file. That's weight theft, and API rate limits are irrelevant to it.
How can teams defend against model extraction and theft?
Protect the files and the API separately. Apply least privilege so callers can get predictions without export rights. Follow API key security practice so every key has an owner and can be revoked on its own. Authentication won't stop a paying customer who's misusing legitimate access.
- Inventory private weights, checkpoints, adapters and backups, and name who needs read or export access.
- Keep model files in private, encrypted storage. Restrict decryption keys as tightly as the bucket.
- Separate training, deployment, inference and export roles, including for CI pipelines and service accounts.
- Issue a distinct credential per customer and workload. Revoke the unused ones.
- Set quotas and rate limits that fit real usage, and review sustained or coordinated collection.
- Return only the output detail the product needs.
- Log artifact access, exports and API calls with account and model version.
- Run an authorized extraction exercise and see whether your controls slow it down and leave usable evidence.
What should teams do when model theft is suspected?
Preserve identity, API, storage and export logs before retention rolls them off. Identify the model versions, accounts and artifacts involved. Revoke compromised credentials and close any exposed storage or export path.
Keep confirmed file access separate from suspected copying. Similar answers alone don't prove extraction. High query volume alone doesn't prove intent. Model owners, security and legal all need to be in the room.
Run it through AI incident response. Revocation stops further access. It won't pull back weights or responses that already left.
How Identra thinks about it
Identra helps protect the credentials and files that lead to proprietary models. Pastes and prompts to supported AI apps are checked on the device for secrets such as API keys and credentials before they're sent, and file uploads to AI can be blocked by policy. Teams can see which connected apps hold admin rights or data access in identity, SaaS and cloud providers, and analysts can revoke risky OAuth grants.
Go deeper: AI security, built on identity
Frequently asked questions
Does model extraction recover the original weights?
Usually not. Query-based extraction tends to produce a substitute that behaves similarly. Getting the actual parameters takes direct weight theft.
Can someone extract a model using legitimate API access?
Yes. A customer can be allowed to request predictions and still be barred from using them to build a copy. Pair access controls with usage review and clear terms.
Are fine-tuned open-weight models at risk?
Yes. The base model being public doesn't make your private adapters or modified weights public. Protect them like any other proprietary artifact.
Do rate limits prevent model extraction?
They slow collection down. They don't stop it, since an attacker can spread queries across accounts or over time.
Is model distillation the same as model theft?
No. Distillation is a normal technique for training one model on another's outputs. It's theft when it's done without permission.
