Model extraction: how to protect your AI gateway
An operational guide to separating legitimate distillation from coordinated abuse, correlating signals, protecting reasoning, and evaluating an AI gateway response.
- Published
- min read
- 8 min read
- Categoría
- AI Systems
On this page
9 chapters- 01Model extraction does not look dangerous in one request
- 02The mechanism: from useful answers to a training corpus
- 031. Bind identity, project, payment, and tenant before the model
- 042. Detect the behavior graph, not a phrase
- 053. Separate the answer from the reasoning artifact
- 064. Apply compound budgets and govern intermediaries
- 075. Run a reversible response and 6. evaluate it with attacks and real customers
- 08Hypothetical example: a B2B copilot behind a router
- 09The decision: protect the complete system, not only the endpoint

Chapter 01
Model extraction does not look dangerous in one request
A prompt asking a model to solve a coding problem, score an answer, or explain a decision may belong to a real customer. The same shape, repeated with small variations across thousands of identities, may also feed a dataset intended to copy another model's capabilities. If the gateway decides from forbidden words alone, it will punish legitimate research and use; if it looks only at volume per API key, a distributed network will split the work and disappear into the average. The detection unit is not an isolated prompt: it is the coordinated pattern that emerges across identities, infrastructure, time, tasks, and outputs.
On September 30, 2026, OpenAI published an investigation into an adversarial distillation campaign. The disclosure appeared four Lima calendar days before this article and is the central news event. OpenAI says it observed spikes of 16,000 requests with a relevant pattern from more than 4,000 users and later connected activity across a cluster exceeding 15,000; its note clarifies that those figures describe attempts, not necessarily successful extractions. It also reports an encrypted-reasoning replay path and controls over streamed output. These are provider findings and attributions, not an independent audit or a benchmark transferable to every product.
Chapter 02
The mechanism: from useful answers to a training corpus
Distillation is a legitimate technique: a teacher model produces signals used to train a smaller or specialized student model. It becomes illicit extraction when someone obtains those capabilities covertly, at scale, and without authorization. A prompt template does not define the boundary. Consent and terms of use matter, as do who controls the data, input provenance, the destination of outputs, and aggregate behavior. A team distilling its own model with an authorized corpus is not equivalent to a network of fraudulent accounts attempting to reconstruct protected capabilities.
In February, Anthropic described campaigns that, according to its investigation, distributed traffic across accounts, payment methods, and access routes while concentrating on reasoning, code, tool use, and evaluation tasks. Its September report adds another risk: third-party routers that retain or relay conversations can expose user data and turn ordinary sessions into training material without the user's knowledge. Both sources are disclosures by the lab itself and should be read with that limitation, but they converge on a useful technical decision: defending the model requires cross-layer correlation; protecting the user also requires data minimization and governance over every intermediary.
Chapter 03
1. Bind identity, project, payment, and tenant before the model
A long API key does not create a trustworthy identity. Issue credentials per application and environment; bind them to organization, tenant, plan, region, and owner; limit their scope; rotate secrets; and record ownership changes. For testing, education, or research, define risk-proportionate verification flows instead of permanent exceptions. If an integrator resells access, require it to preserve customer identity and abuse signals, because an unattributed aggregator turns thousands of actors into one blind spot.
The response to an anomaly should degrade gracefully. Instead of jumping from full access to permanent blocking, prepare tighter limits, additional review, cooldowns, suspension of a specific capability, and human escalation. Preserve an appeal path and enough evidence to explain the decision without disclosing rules that make evasion easier. The goal is not to guess every user's moral intent, but to reduce the ability to coordinate abuse while preventing a distributed campaign from immediately reappearing under another key.
Chapter 04
2. Detect the behavior graph, not a phrase
Build signals per request and per cohort: cadence, structural similarity, true task diversity, input-to-output ratio, selected models, retries, errors, regions, ASN, device, payment method, and synchronization across accounts. You do not need to retain every full conversation to preserve all of those signals. Derive features with bounded retention, hashes resistant to trivial leakage, time windows, and separate access controls. Thresholds must learn from each product's legitimate traffic; a nightly evaluation batch may look coordinated and be fully authorized.
Correlate across multiple horizons. A ten-minute burst can reveal automation; weekly repetition can reveal account replenishment; a sharp change after a model launch can show adaptation. Even then, correlation is a signal, not a verdict. Combine deterministic rules for clear limits with models or statistical analysis for clusters, and record which evidence triggered each action. If the detector works only through an analyst's memory, it cannot be evaluated or recovered when the adversary changes.
Chapter 05
3. Separate the answer from the reasoning artifact
Do not design a contract that depends on revealing private internal reasoning. Ask for verifiable outputs: the answer, citations, summarized calculations, tool calls, decisions, and evidence needed for review. Treat any opaque or encrypted state traveling between turns as a capability token: bind it to user, workspace, model, purpose, and expiry; prevent another tenant from replaying it; and do not write it into logs, analytics, or support tools as if it were harmless text. Portability without context can turn a continuity optimization into a replay surface.
Streaming output also needs a control point. If every token goes directly to the browser, you cannot hold a sequence detected late without accepting that some of it has already left. Define small buffers for high-risk classes, size limits, leakage detectors, and a safe response when the stream is interrupted. Do not present this filter as infallible: it can create false positives or degrade latency. Measure both effects and let the product choose which routes require stronger inspection.
Chapter 06
4. Apply compound budgets and govern intermediaries
A requests-per-minute limit is necessary but insufficient. Budget tokens, concurrency, output volume, cost, high-capability tasks, and account expansion across user, organization, and shared infrastructure. Add dynamic limits by risk and model novelty: a valuable reasoning route may deserve different controls than a short classification. Prevent customers from multiplying quota through new projects, invitations, or keys; the calculation must roll up to an identity and contract rather than stopping at the easiest identifier to replace.
Map every path to the provider: first-party backend, partner cloud, fallback, router, SDK, observability, and support. The same control must follow a request when it crosses a partner, or an attacker will choose the route with the least visibility. Contractually require processing purpose, retention, training terms, subprocessors, region, incident response, and signal return. Test those promises technically. An architecture diagram that omits the commercial router also omits who can read the conversation and who must detect coordinated abuse.
Chapter 07
5. Run a reversible response and 6. evaluate it with attacks and real customers
Define incident states before you need them: observe, limit, hold output, require reverification, suspend, investigate, and close. Each state needs an owner, minimum evidence, SLA, scope, and exit criterion. Preserve samples and features with a chain of custody and privacy controls; notify partners when their route is involved; and maintain a list of downstream consumers so contaminated data or compromised credentials can be withdrawn. External coordination does not replace local control, but it prevents the same cluster from moving between providers without friction.
The evaluation should mix synthetic attacks, replay of closed incidents, and difficult legitimate traffic: evaluation teams, batch workloads, classrooms, CI, and customers using repeated templates. Measure campaign-level recall, action-level precision, time to correlation, new accounts after enforcement, leakage before blocking, added latency, and reversed appeals. Separate detection from containment: a system can alert correctly and block too late. Keep a holdout set, refresh cases as models and routes change, and run exercises where one critical signal disappears to verify that the defense degrades safely.
Chapter 08
Hypothetical example: a B2B copilot behind a router
This example is hypothetical. A support copilot summarizes tickets, drafts replies, and evaluates quality. Over one week, traffic grows across hundreds of new accounts. Every account respects its quota and its prompts look normal, but they share infrastructure, rotate through the same technical domains, request long outputs, and run nearly identical rubrics in synchronized windows. The gateway groups signals without retaining full attachments, reduces the quota for evaluation tasks, and requires reverification. The team reviews authorized samples before suspending the cluster.
The investigation finds that some traffic passed through a router that replaced end-user identities with one shared credential. The team cannot attribute every request, so it temporarily disables the high-capability route, preserves short classification for affected customers, and asks the partner for tenant-level signals. It also rotates leaked keys and notifies users whose data may have been relayed. The lesson is not that the detector “found a bad prompt”: it found a combined identity, observability, contract, and budget failure, then responded without shutting down the entire product.
Chapter 09
The decision: protect the complete system, not only the endpoint
The recent disclosures do not prove that every large-scale automation is extraction or that one defense generalizes across providers. They do show why isolated controls fail: identity can fragment, traffic can move through partners, an output can reveal more than intended, and a campaign can adapt after enforcement. Start with six verifiable controls: bound identity, a behavior graph, reasoning separation, compound budgets, reversible response, and continuous evaluation. Also document what you cannot observe and which residual risk you accept.
A secure gateway is not a proxy with rate limiting. It is the layer that preserves identity context, applies policy, minimizes data, observes patterns, governs providers, and produces evidence for action. If you need to design that layer inside a B2B agent or copilot, Wasyra's custom agent service can help turn these controls into contracts, traces, and evaluations. It is an implementation service separate from the product at agents.wasyra.com; the principle remains the same: no capability reaches production without a boundary that can be measured, explained, and stopped.
Written by
Wasyra AI Systems
Trust, copilots, and enterprise adoption
Wasyra AI Systems covers guardrails, suggestion-first modes, and review design so work assistants earn real adoption.
Series
AI systems that actually reach production
A series on agents, copilots, and guardrails for bringing AI into real work without breaking trust or operations.
Posts in this seriesMore from this author
More from this author
AI Systems
Async agents: how to keep work moving without blocking
Five gates that let an agent advance while slow tools run without losing dependencies, control, state, or evidence.
ArticleAI Systems
Computer-use agents: how to evaluate desktop automation
Five gates for deciding whether an agent that sees, clicks, and types in desktop apps can operate without losing control or evidence.
ArticleKeep reading
Keep reading
AI Systems
Async agents: how to keep work moving without blocking
Five gates that let an agent advance while slow tools run without losing dependencies, control, state, or evidence.
ArticleAI Systems
Computer-use agents: how to evaluate desktop automation
Five gates for deciding whether an agent that sees, clicks, and types in desktop apps can operate without losing control or evidence.
ArticleAI Systems
HydraFusion: how to evaluate multi-model orchestration
Single, cascade, or critique are not magic shortcuts. Five gates measure quality, cost, latency, isolation, and repository state.
Article