Source
Xiaonan Xu, Wenjing Wu, "MCP Error Messages Written for Developers Hurt the Most Capable Agents Most", arXiv:2609.35381 [cs.SE; cs.AI], submitted 2026-09-28 (v1), revised 09-29 (v2 fixes quotation marks only — numbers unchanged), 15 pages, 6 tables, submitted to the Journal of Systems and Software. Affiliations: College of Computing, Georgia Institute of Technology (corresponding author), and Department of Computer Science, University of Colorado Boulder. Figures were checked against the arXiv HTML full text (v1 and v2) on 2026-10-10.
This is a preprint under journal review. No funding is stated and the authors declare no competing interests. Scenarios, error texts, survey data, model outputs, and code are public on GitHub (WenJing95/tool-error-text), so the work can be replicated. All five tested models are OpenAI models, however, so whether the result holds for other vendors' models is untested.
Study overview
Many MCP servers wrap web APIs built for human developers (in prior work, 88.6% of 116 official servers are REST-backed). Their error messages therefore suggest things a person can do: "run this command in a terminal," "change the setting," "wait and try again." What happens when the reader is an agent that can only call tools, like an MCP connector in a chat app? The paper answers in three steps.
First, it located 3,001 error messages in the source of the 150 most-starred MCP servers on GitHub (official SDK, updated within the past year, 426–186,635 stars) and classified them (labels assigned by OpenAI Codex agents following a published codebook). Second, it built 168 scenarios — 24 for each of seven failure types — from BFCL V4 multi-turn tasks and ran five models (GPT-5.5, GPT-5.6 Sol, GPT-6 Sol, GPT-6 Astra, GPT-6 Luna) across six error texts in 15,120 trials (reasoning effort high, at most eight tool calls, accessed 24, 25 and 27 September 2026). Third, it tested two remedies: server authors rewriting the step as a call to a server tool, and agent developers deleting the step with a one-sentence prompt (1,440 additional filter trials).
Key results
In the survey, 949 of 3,001 messages stated a next step, and 477 of those depended on caller conditions the server cannot see (tools, login, permissions). Among 209 credential, permission, and rate-limit messages, 99 of 128 steps were caller-dependent, and 93 of those asked for something an MCP-only agent cannot do. On credential errors, 62 of 67 steps asked for a configuration change, a web page, or a terminal command; on rate limits, 20 of 30 steps said to wait and retry, and none named the call to repeat.
- 55Config change · 59.1%
- 24Action on a web page · 25.8%
- 13Terminal command · 14.0%
- 1Wait · 1.1%
Source: paper Section 3 (text of Table 1)
View as table
| Item | Count |
|---|---|
| Config change | 55 |
| Action on a web page | 24 |
| Terminal command | 13 |
| Wait | 1 |
In the experiment, the agents did what the text said. With an expired session and only the cause stated, 82% recovered on average by calling the login tool themselves. Appending the single sentence "Please run: reddit-mcp-buddy --auth" dropped recovery to 45%, and 55% of trials ended without a repair. Rewriting the step as a server tool call ("Call ticket_login first.") restored it to 84%, and deleting the step with a one-sentence prompt (run on GPT-6 Luna) restored it to 82%.
- Original (terminal command)
- Cause alone
- Rewritten (login tool)
- Step deleted by prompt
Source: paper Table 4 (arXiv:2609.35381v2)
View as table
| Item | Original (terminal command) | Cause alone | Rewritten (login tool) | Step deleted by prompt |
|---|---|---|---|---|
| GPT-5.5 | 58% | 76% | 85% | 86% |
| GPT-5.6 Sol | 57% | 92% | 89% | 89% |
| GPT-6 Sol | 46% | 85% | 82% | 76% |
| GPT-6 Astra | 6% | 75% | 75% | 75% |
| GPT-6 Luna | 57% | 83% | 88% | 85% |
| Mean of five | 45% | 82% | 84% | 82% |
The loss grew with newer, larger models: 18 points for GPT-5.5, 35 for GPT-5.6 Sol, 39 for GPT-6 Sol, and 69 for the largest, GPT-6 Astra (26 for the small GPT-6 Luna). The 51-point gap between Astra and GPT-5.5 has a 95% interval of 29–72, clear of zero. Under the command, Astra made only 0.60 tool calls per trial, and in 48 of its 68 trials without a repair it handed the repair back to the user ("please reconnect"), even though the login tool was in its tool list.
Rate limits are starker still. Under GitHub's "Wait before retrying.", mean recovery was 6% and 89–99% of trials simply ended. Naming the call to repeat raised recovery to 88%, a gain of 82 points (interval 75–88), at the cost of about one extra call: tool calls rose from 0.37 to 1.42 and tokens from 3,098 to 5,409 per trial.
- Original (bare wait)
- Names the call to repeat
Source: paper Table 5
View as table
| Item | Original (bare wait) | Names the call to repeat |
|---|---|---|
| GPT-5.5 | 4% | 78% |
| GPT-5.6 Sol | 7% | 99% |
| GPT-6 Sol | 8% | 96% |
| GPT-6 Astra | 11% | 89% |
| GPT-6 Luna | 1% | 79% |
| Mean of five | 6% | 88% |
When the failure is visible in the agent's own call (wrong unit or format, missing field, wrong tool, missing resource), every text stayed within four points of the generic notice (Table 6). The text decided the outcome only for failures visible solely through the error message — expired sessions, missing permissions, rate limits. On a missing permission, the cause alone gave 0% recovery and naming the login tool gave 53%.
Credibility assessment
Reasons to trust it: comparisons are within scenario, changing only the text from the same saved state, so no other variable leaks in; recovery is judged by BFCL's state check rather than an LLM judge, removing scoring subjectivity. Intervals come from 10,000 bootstrap resamples of scenarios within failure type, all data and code are public, and the experimental texts reuse wording found in real servers.
Caveats: (1) All five models are OpenAI's; Claude, Gemini, and others are absent. (2) The interpretation that newer models follow text more literally is consistent with OpenAI and Anthropic prompting guides but is not causally tested. (3) Survey labels were assigned by Codex agents, not humans, and no human agreement rate is reported. (4) In the BFCL setup, a retried rate-limited call succeeds immediately (no wait is enforced), so whether "call again" works as well against real services is a separate question.
Related work: MCP ecosystem studies pointing the same way
- Mastouri et al. (2025). From REST to MCP: An Empirical Study of API Wrapping and Automated Server Generation for LLM Agents — the source of this paper's premise: 88.6% of 116 official MCP servers are REST-backed and 92% implement tools as bare API wrappers. It is the structural reason developer-facing text flows straight to agents.
- Hasan et al. (2026). Model Context Protocol (MCP) Tool Descriptions Are Smelly! — the same lens applied to tool descriptions. This paper moves it to error text, showing that failure responses, not just descriptions, are an interface that shapes agent performance.
- Liu et al. (2026). When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks — independent evidence in the same direction: the most capable models most reliably follow instructions planted in tool descriptions. Malicious instructions there, good-faith ones here, but the conclusion overlaps — newer models treat text returned by tools as instructions.
| Angle | This paper | Prior work | Relationship |
|---|---|---|---|
| Where the text comes from | Error text from good-faith server authors | Tool descriptions planted by attackers (2605.24069) | Both read as "tool text = instruction" |
| Capability vs. vulnerability | Bigger loss for newer, larger models (18→69 pts) | Most capable models follow most faithfully (2605.24069) | Same direction |
| Root of the problem | Inherited developer-facing API text | 88.6% REST wrapping (2507.16044) | Supplies the cause |
Prior work showed separately that tool descriptions are defective, that most MCP servers are API wrappers, and that strong models follow planted instructions. This paper measures, in a controlled experiment, the point where all three meet: strong models take the developer-facing error text that wrappers inherit as an instruction, and stop.
Reviewer's judgment
First, we judge the most important number here to be not 69 points but the 82% under "cause alone." The agents could fix the problem all along. Failure came not from model capability but from one helpful-sounding sentence, which shows that the intuition "more guidance = better results" can be wrong in prompt and tool design.
Second, the fact that upgrading the model makes the same server text more costly matters operationally. An agent that works today could see its expired-credential recovery collapse from a model version upgrade alone, so model-swap regression tests must include error-path scenarios. An eval set that covers only success paths will not see this regression.
Third, the two remedies sit in different places and should be kept apart in practice. If you run the server, rewriting the text as a server tool call is the surest fix (84% on credentials, 88% on rate limits). If you connect other people's servers, a step-deletion filter is cheap (0.09 USD to process all 949 messages) and effective. But on missing permissions and rate limits the cause alone barely recovers (0% and 5%), so applying the filter across the board would delete correct steps too — the authors themselves recommend limiting it to credential errors.
Fourth, this connects to the semantic-layer debate. What an agent can do (its tool list) and what it is told to do on failure (error text) need to be aligned in the same vocabulary, so we see error text as part of the tool ontology that should be designed deliberately.
Practical levers
- Name server tools in error text — instead of "please re-authorize," write "call auth_login first"; instead of "wait before retrying," write "wait a few seconds and call place_order again." It is the only form that also works for human developers.
- Move human guidance to a separate field — put terminal commands, config paths, and web links in a field shown to human operators, and leave only the cause and tool calls in the text the model reads (the same idea as RFC 9457 problem details).
- Filter third-party servers on credential errors only — insert a one-sentence prompt that deletes "next step" sentences only when a tool response is a credential error. Don't apply it to permission or rate-limit errors.
- Add error paths to model-upgrade regression sets — keep expired-session, missing-permission, and rate-limit scenarios per failure type, and track the "ended without repair" and "handed back to user" rates.
- Audit error text before connecting — before adding a new MCP server, check its source or sample responses for terminal, config, or web instructions. By this survey's count, 93% of credential-error steps (62 of 67) contained one.
Conclusion
This paper shows with public data that an MCP agent's failure can come from a single sentence the server returns rather than from the model, and that the cost can grow with newer models. The results are limited to OpenAI models and the explanation of model behavior is still a hypothesis, but the fixes are simple and the measured effects large, so both server and agent developers can apply them right away.
What operations teams should check as agents connect to more MCP servers and semantic layers continues in our blog post from the same day, Google's Gemini Agent Connects to Any MCP Server and a Knowledge Catalog.
References
- MCP Error Messages Written for Developers Hurt the Most Capable Agents Most — arXiv abstract
- Same paper, HTML full text (v2) — used to check Tables 1, 4, 5, 6
- Authors' public data and code (GitHub)
- From REST to MCP (arXiv:2507.16044)
- MCP Tool Descriptions Are Smelly! (arXiv:2602.14878)
- When the Manual Lies: MCP Poisoning Benchmark (arXiv:2605.24069)
- Google's Gemini Agent Connects to Any MCP Server and a Knowledge Catalog — sunny34.com blog
Ask AI about this review
The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…