Splitting the Surface Flips the Ranking

COPEX ran 3,375 trials across nine models and measured an average attack success rate (ASR) of 64.4%. Broken out by surface, the numbers range from 58.3% (model/agent) to 71.4% (transport), and Opus-4.7 — the model with the lowest overall ASR at 44.3% — jumps to 73.3% on transport attacks alone, above the overall average.

The operational takeaway is straightforward: a single aggregate ASR in a model-selection report hides exactly the surface your deployment is actually exposed to. Map which surface — client, server, or transport — your MCP deployment touches most before comparing models on one number.

Attacks the Model Never Sees

MCP rebinding — a zero-TTL DNS trick that redirects a connection to an attacker's server — held at 100% ASR across all nine models and across all four defense configurations tested (no scanning, input scanner only, context scanner only, both combined). It happens below the point the model or any scanner can observe: the user task and the tool schema and response.

Combining input and context scanners did cut the average ASR for the other eight attack types from 81.6% to 41.1%. That improvement never reaches the transport layer, though — without network-level controls such as certificate pinning, DNS response validation, or mTLS, the same number repeats indefinitely.

The Judging Criteria Change the Numbers

Scoring the same six attacks with a deterministic trace judge yields an average of 77.1%; switching to an LLM (semantic) judge pushes it to 83.3%, a 6.3-point gap. Swapping the reasoning scaffold from ReAct to Reflexion drops average ASR from 86.2% to 82.5%, but the direction reverses by model — Sonnet-4.6 falls 9.9 points while Grok-4.3 rises 0.5.

Reporting an ASR externally, or comparing it across vendors, without stating the judging method means the same attack reads as a different number. And since a reasoning-scaffold swap like Reflexion helps some models and hurts others, it cannot be assumed to be a reliable defense on its own.

From Design to Operations: An MCP Red-Team and Transport Gate Checklist

Set combined input-and-context scanning to cut average ASR by at least 49.6% relative as the first target, and track the transport layer with network metrics rather than model metrics: 100% certificate-pin coverage and zero DNS-validation failures a week are reasonable acceptance bars.

Make surface-level ASR a required field in every model evaluation report. Selecting on a single overall average would miss cases like Opus-4.7, whose transport-surface number reverses its overall ranking.

The four most common mistakes: picking a model from its overall average alone, treating a passed scanner as "secure," trusting a reasoning-scaffold change like Reflexion as a security control, and comparing ASR figures without disclosing the judging method.

When a DNS response with TTL=0 or a certificate outside the registered pin set appears, cut the connection immediately and route to human approval instead of retrying automatically — an automatic retry just hands the attacker the same window again.

Before shipping, reproduce COPEX's 25 attack types across its four surfaces against your own MCP server inventory. State the judging method — deterministic trace or LLM judge — in every report, and standardize log fields for surface tag, attack type, judging method, and certificate-pin status. Mask PII that leaks into tool-response logs through the same pipeline.

Track surface-level ASR on a weekly dashboard, and add newly published attack types to the red-team set as they appear in academic or vendor disclosures. Keep model-swap changelog entries separate from network-control changelog entries, so that when ASR moves next quarter you can tell whether the model or the infrastructure caused it.

Takeaways at a Glance

MCP agent security does not finish with a model swap or a prompt tweak. Put surface-level ASR into model-selection criteria, handle the transport layer with network controls like certificate pinning and mTLS, and lock a red-team reproduction with a stated judging method into the pre-deploy checklist — only then do these numbers become operationally useful.

References

COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers — arXiv

Attack Success Rate 64.4%: COPEX Splits Model Blame From Lower-Layer MCP Compromise — sunny34.com Research

Ask AI about this article

The assistant has read this article. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…