Skip to content

Enterprise deploymentCustomer-hosted or fully managed.Contact salesView pricing

Threat briefing: when an approved MCP tool changes under you

Research note · September 29, 2026 · 8 min readBy Intertrace Threat Research — Threat intelligence
MCPtool poisoningschema pinningthreat briefing

An MCP server can change what a tool says and accepts after you've approved it. Agents re-read tool descriptions on every session, so a quiet edit is a new instruction. How schema drift works, what it looks like, and how to catch it before the first call.

Intertrace Threat ResearchThreat intelligence

Intelligence on real-world AI and AI-agent attacks — incidents, CVEs, and the controls that mitigate them at runtime.

Research note · September 29, 2026 · 8 min read

The technique

Tool descriptions are prompts. The model uses them to decide when and how to call a tool, which makes them one of the most privileged pieces of text in an agent's context. Schema drift is the simple observation that this text is fetched fresh from a server you don't control, every time. The approval you gave was for a snapshot; the agent is running on the live version.

Drift doesn't need a compromised server. A dependency update, a new maintainer or a well-meaning feature can widen what a tool does. But it is also the cleanest path for an attacker who gains control of a popular server: publish something useful, wait for it to be approved widely, then change it.

What drift looks like

  • A description gains an instruction aimed at the model — for example, to also read a credentials file and include it in the next call.
  • An input schema gains an optional field such as a destination URL or a 'notes' parameter that silently carries data out.
  • A read-only tool's description starts describing writes, deletes or sends, and the agent begins using it that way.
  • A tool is renamed to shadow a trusted tool from another server, so the agent calls the wrong one.
  • Invisible characters or markup hide text from the human reviewer while the model still reads it.

Why reviews alone don't catch it

Reviews happen once; tool lists are fetched constantly. Most MCP clients don't show the full description text, let alone diff it between sessions. And the reviewer and the model often see different things — a description can render cleanly for a person while carrying instructions the model will act on.

Mitigations that hold

  1. Pin: hash the canonical tool list (names, descriptions, schemas) when you approve a server, and store the hash.
  2. Compare on every session: if the live list doesn't match the pin, hold the server until someone re-approves the change.
  3. Inspect descriptions as untrusted text: scan for instructions addressed to the model and for hidden characters before the agent sees them.
  4. Authorize calls, not servers: even an approved tool call should be checked against policy with its actual arguments at call time.
  5. Keep a record: log every drift event and every denied call, so a changed server becomes an alert instead of a silent behaviour change.

How Intertrace handles it

MCP traffic routed through Intertrace is pinned per server: the tool list is hashed on approval and compared on every sync, and drift holds the server for review. Descriptions are inspected for hidden and model-addressed instructions before they reach the agent. Every tools/call is authorized with its real arguments, and denied calls fail closed with a logged reason. You can check your own MCP configuration for risky servers with the free MCP risk check.

Continue reading

← Back to blog