Source

Nahom Birhan, Mehrdad Rostamzadeh, Sidhant Narula, Mahmoud Nazzal, Mohammad Ghasemigol, Daniel Takabi, "COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers", arXiv:2610.04378 [cs.CR], submitted 2026-10-03, DOI 10.48550/arXiv.2610.04378. Affiliation: Old Dominion University. Submitted to the Third Workshop on "Agents in the Wild: Safety, Security, and Beyond." Benchmark and evaluation code released (github.com/inspire-center/copex). The full text was cross-checked against a primary-source snapshot fetched by fetch_source_snapshot.py at 2026-10-06T22:10:52Z (this session's network egress was blocked).

This is an unreviewed preprint. It is marked as a workshop submission, and workshop review is typically lighter than a main-conference track, with no published acceptance rate — treat it as close to a first draft. All authors are affiliated with Old Dominion University, and no funding source or conflict-of-interest statement is given. There is also no indication of support from the vendors behind the evaluated models (Anthropic, OpenAI, Google, xAI), which reduces conflict-of-interest concerns for a third-party academic evaluation.

Study Overview

The question is simple: in an MCP-enabled agent, when an attack succeeds, can you tell whether the model is at fault, or whether the fault lies in a lower layer (client, server, or transport) that the model never observes? Prior benchmarks mostly evaluate deployed agent products, which conflates model susceptibility with the product's guardrails, orchestration, and general task capability. The authors instead fix the agent loop, orchestration, tool environment, and execution budget and vary only the tool-selecting LLM, producing a controlled environment called COPEX. It wires up 16 MCP servers exposing 75 tools and 7 resources, places 25 attack types across four entry surfaces (model/agent, client, server/tool, transport), and runs five scenario variants per attack type for 125 total scenarios.

Key Results

Across nine models (Opus 4.7, Sonnet 4.6, GPT-5.5, GPT-5 Mini, GPT-4.1, Gemini 3.1, Grok 4.3, Llama-3.2:3b, Qwen-2.5:3b) and 3,375 trials, the main benchmark's mean attack success rate (ASR) was 64.4%. Surface-level means ranged from 58.3% (model/agent) to 71.4% (transport) — attacks that never touch the user prompt succeeded as often or more.

ModelOverall mean ASRClient (cl)Transport (tr)
Opus-4.744.3%38.1%73.3%
Sonnet-4.645.1%37.1%62.2%
Gemini-3.174.1%95.2%75.6%
Grok-4.381.6%98.1%80.0%
Mean (9 models)64.4%67.7%71.4%

Opus-4.7 has the lowest overall ASR, yet jumps to 73.3% on transport attacks — no model stays safe across every surface. MCP rebinding — a transport-surface attack that uses zero-TTL DNS responses to redirect the connection to an attacker-controlled endpoint — happens below the model's observation boundary, so it hit 100% ASR for all nine models.

Defense effectiveness has to be read together with the outcome oracle. On an eight-attack defense subset, combining input and context scanning dropped mean ASR from 81.6% to 41.1% (a 40.5pp reduction, 49.6% relative). But in the same experiment, MCP rebinding stayed at 100% under all four defense configurations — it happens below what the scanners can observe (the user task, tool schemas, and tool responses), so no amount of context scanning can catch it.

Defense modeMean ASRMCP rebinding ASR
No scanning81.6%100%
Input scanner only48.4%100%
Context scanner only57.6%100%
Input + context combined41.1%100%

Switching the outcome oracle also moves the number. Scoring the same six attacks with deterministic trace predicates gives a mean of 77.1%; switching to an LLM judge (semantic) raises it to 83.3% (+6.3pp). Swapping ReAct for Reflexion self-critique drops mean ASR from 86.2% to 82.5% (−3.7pp), but the direction is model-dependent — from −9.9pp for Sonnet-4.6 to +0.5pp for Grok-4.3 — so it's not a consistent safeguard.

Credibility Assessment

Reasons to trust it: the controlled design holds the agent stack fixed and varies only the model, so deployment-product guardrail and orchestration effects don't get mixed in with the model's own susceptibility. The independently verified prior MCP and agent-security benchmarks below report attack success rates of similar magnitude, which supports that this result isn't an outlier.

Caveats: this is a workshop submission with light peer review and no stated conflict-of-interest disclosure. The defense and reasoning-architecture studies ran only one trial per variant, which the authors themselves flag as "point estimates" rather than precise measurements. 15 of the 25 attack types rely on an LLM judge (Claude Haiku 4.5), so their reported ASR reflects that judge's decision boundary. ASR also only measures susceptibility once an attack is delivered — it says nothing about how often such an attack would actually be attempted or how accessible the attack surface is in practice.

Related Work

  • Debenedetti et al. (2024), AgentDojo — Prior work (NeurIPS 2024 Datasets and Benchmarks track). Found that across 97 realistic tasks, LLMs solve less than 66% even with no attack present, supporting the same premise COPEX starts from: task capability and security susceptibility need to be measured separately.
  • Zhang et al. (2025), Agent Security Bench (ASB) — Prior work (accepted at ICLR 2025). Widens the attack surface to system prompts and memory across 10 scenarios and reports a top mean ASR of 84.30%. Broader in scope than COPEX, but it does not separate the MCP-specific client and transport layers — a contrast this paper makes explicit.
  • Yang, Wu, and Chen (2025), MCPSecBench — Prior work this paper directly compares against in its own Table 1. Applies 17 attack types to deployed products including Claude Desktop and Cursor and reports that over 85% of attacks compromised at least one platform, illustrating the difference between "evaluating a deployed product" and "isolating the model."

All three prior benchmarks report LLM-agent attack success rates in a similar, roughly-majority range, so this paper's 64.4% mean is not an outlier reading but a consistent observation within the same literature. What's distinct about COPEX is that it attributes that success rate across the model, client, server, and transport layers rather than reporting a single number.

Reviewer's Take

First, I'd argue the most practically important conclusion here isn't the 64.4% average but the fact that an unsplit ranking lies. Opus-4.7 has the lowest overall ASR, yet scores 73.3% on transport attacks — above the 64.4% average. Picking a model off a single overall ASR number misses the risk on whichever surface your actual deployment is exposed through.

Second, I think the paper underweights its own most striking finding: MCP rebinding stays at exactly 100% across every defense mode, every one of the nine models, and both reasoning architectures. That means investing in the model or the prompt cannot reduce this attack at all — only transport-layer controls (certificate pinning, DNS response validation, mTLS) can. The authors frame this as being "outside the observation boundary" by design rather than calling it out as a limitation, which understates how material it is operationally.

Practical Takeaways

  • Report ASR by surface, not just overall — when evaluating a model for MCP deployment, ask for model/client/server/transport breakdowns and check which surfaces your own deployment actually exposes.
  • The transport layer can't be fixed by the model — DNS rebinding and endpoint hijacking need network-level controls (certificate pinning, mTLS), not a better model or prompt.
  • Combine input and context scanning — the combination outperforms either scanner alone (81.6%→41.1%), though attacks below the scanners' observation boundary still pass through.
  • Disclose the outcome oracle — state whether a reported ASR uses deterministic trace checks or an LLM judge; the same six attacks differed by +6.3pp between the two.
  • Don't treat Reflexion as a security control — adding self-critique moved ASR in opposite directions across models (−9.9pp to +0.5pp), so it isn't a consistent defense.

Conclusion

COPEX's contribution isn't a new attack technique — it's a measurement design that assigns blame. A mean attack success rate of 64.4%, a 49.6% relative drop from combined scanning, and yet 100% ASR on transport attacks under every defense tested together show, in numbers, that swapping models or tuning prompts alone cannot secure an MCP-enabled agent. Because this is a workshop submission with light review and the defense/architecture studies are point estimates, teams should re-verify these surface-level numbers on their own stack before relying on them. The operational follow-through for vetting third-party MCP servers is covered in MCP Server Trust and Version Gates.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…