Source

Xiaonan Xu, Wenjing Wu, "MCP Error Messages Written for Developers Hurt the Most Capable Agents Most", arXiv:2609.35381 [cs.SE; cs.AI], submitted 2026-09-28 (v1), revised 09-29 (v2 fixes quotation marks only — numbers unchanged), 15 pages, 6 tables, submitted to the Journal of Systems and Software. Affiliations: College of Computing, Georgia Institute of Technology (corresponding author), and Department of Computer Science, University of Colorado Boulder. Figures were checked against the arXiv HTML full text (v1 and v2) on 2026-10-10.

This is a preprint under journal review. No funding is stated and the authors declare no competing interests. Scenarios, error texts, survey data, model outputs, and code are public on GitHub (WenJing95/tool-error-text), so the work can be replicated. All five tested models are OpenAI models, however, so whether the result holds for other vendors' models is untested.

Study overview

Many MCP servers wrap web APIs built for human developers (in prior work, 88.6% of 116 official servers are REST-backed). Their error messages therefore suggest things a person can do: "run this command in a terminal," "change the setting," "wait and try again." What happens when the reader is an agent that can only call tools, like an MCP connector in a chat app? The paper answers in three steps.

First, it located 3,001 error messages in the source of the 150 most-starred MCP servers on GitHub (official SDK, updated within the past year, 426–186,635 stars) and classified them (labels assigned by OpenAI Codex agents following a published codebook). Second, it built 168 scenarios — 24 for each of seven failure types — from BFCL V4 multi-turn tasks and ran five models (GPT-5.5, GPT-5.6 Sol, GPT-6 Sol, GPT-6 Astra, GPT-6 Luna) across six error texts in 15,120 trials (reasoning effort high, at most eight tool calls, accessed 24, 25 and 27 September 2026). Third, it tested two remedies: server authors rewriting the step as a call to a server tool, and agent developers deleting the step with a one-sentence prompt (1,440 additional filter trials).

Key results

In the survey, 949 of 3,001 messages stated a next step, and 477 of those depended on caller conditions the server cannot see (tools, login, permissions). Among 209 credential, permission, and rate-limit messages, 99 of 128 steps were caller-dependent, and 93 of those asked for something an MCP-only agent cannot do. On credential errors, 62 of 67 steps asked for a configuration change, a web page, or a terminal command; on rate limits, 20 of 30 steps said to wait and retry, and none named the call to repeat.

What the 93 out-of-reach next steps ask forCredential, permission and rate-limit messages in 150 popular MCP servers
  • 55Config change · 59.1%
  • 24Action on a web page · 25.8%
  • 13Terminal command · 14.0%
  • 1Wait · 1.1%

Source: paper Section 3 (text of Table 1)

View as table
ItemCount
Config change55
Action on a web page24
Terminal command13
Wait1

In the experiment, the agents did what the text said. With an expired session and only the cause stated, 82% recovered on average by calling the login tool themselves. Appending the single sentence "Please run: reddit-mcp-buddy --auth" dropped recovery to 45%, and 55% of trials ended without a repair. Rewriting the step as a server tool call ("Call ticket_login first.") restored it to 84%, and deleting the step with a one-sentence prompt (run on GPT-6 Luna) restored it to 82%.

Recovery on expired credentials (%)4 error texts × 5 models, mean of 24 scenarios × 3 runs
  • Original (terminal command)
  • Cause alone
  • Rewritten (login tool)
  • Step deleted by prompt
GPT-5.5
58%
76%
85%
86%
GPT-5.6 Sol
57%
92%
89%
89%
GPT-6 Sol
46%
85%
82%
76%
GPT-6 Astra
6%
75%
75%
75%
GPT-6 Luna
57%
83%
88%
85%
Mean of five
45%
82%
84%
82%

Source: paper Table 4 (arXiv:2609.35381v2)

View as table
ItemOriginal (terminal command)Cause aloneRewritten (login tool)Step deleted by prompt
GPT-5.558%76%85%86%
GPT-5.6 Sol57%92%89%89%
GPT-6 Sol46%85%82%76%
GPT-6 Astra6%75%75%75%
GPT-6 Luna57%83%88%85%
Mean of five45%82%84%82%

The loss grew with newer, larger models: 18 points for GPT-5.5, 35 for GPT-5.6 Sol, 39 for GPT-6 Sol, and 69 for the largest, GPT-6 Astra (26 for the small GPT-6 Luna). The 51-point gap between Astra and GPT-5.5 has a 95% interval of 29–72, clear of zero. Under the command, Astra made only 0.60 tool calls per trial, and in 48 of its 68 trials without a repair it handed the repair back to the user ("please reconnect"), even though the login tool was in its tool list.

Rate limits are starker still. Under GitHub's "Wait before retrying.", mean recovery was 6% and 89–99% of trials simply ended. Naming the call to repeat raised recovery to 88%, a gain of 82 points (interval 75–88), at the cost of about one extra call: tool calls rose from 0.37 to 1.42 and tokens from 3,098 to 5,409 per trial.

Recovery on rate limits (%)"Wait before retrying." vs. "Wait a few seconds and call <tool> again"
  • Original (bare wait)
  • Names the call to repeat
GPT-5.5
4%
78%
GPT-5.6 Sol
7%
99%
GPT-6 Sol
8%
96%
GPT-6 Astra
11%
89%
GPT-6 Luna
1%
79%
Mean of five
6%
88%

Source: paper Table 5

View as table
ItemOriginal (bare wait)Names the call to repeat
GPT-5.54%78%
GPT-5.6 Sol7%99%
GPT-6 Sol8%96%
GPT-6 Astra11%89%
GPT-6 Luna1%79%
Mean of five6%88%

When the failure is visible in the agent's own call (wrong unit or format, missing field, wrong tool, missing resource), every text stayed within four points of the generic notice (Table 6). The text decided the outcome only for failures visible solely through the error message — expired sessions, missing permissions, rate limits. On a missing permission, the cause alone gave 0% recovery and naming the login tool gave 53%.

Credibility assessment

Reasons to trust it: comparisons are within scenario, changing only the text from the same saved state, so no other variable leaks in; recovery is judged by BFCL's state check rather than an LLM judge, removing scoring subjectivity. Intervals come from 10,000 bootstrap resamples of scenarios within failure type, all data and code are public, and the experimental texts reuse wording found in real servers.

Caveats: (1) All five models are OpenAI's; Claude, Gemini, and others are absent. (2) The interpretation that newer models follow text more literally is consistent with OpenAI and Anthropic prompting guides but is not causally tested. (3) Survey labels were assigned by Codex agents, not humans, and no human agreement rate is reported. (4) In the BFCL setup, a retried rate-limited call succeeds immediately (no wait is enforced), so whether "call again" works as well against real services is a separate question.

Related work: MCP ecosystem studies pointing the same way

AngleThis paperPrior workRelationship
Where the text comes fromError text from good-faith server authorsTool descriptions planted by attackers (2605.24069)Both read as "tool text = instruction"
Capability vs. vulnerabilityBigger loss for newer, larger models (18→69 pts)Most capable models follow most faithfully (2605.24069)Same direction
Root of the problemInherited developer-facing API text88.6% REST wrapping (2507.16044)Supplies the cause

Prior work showed separately that tool descriptions are defective, that most MCP servers are API wrappers, and that strong models follow planted instructions. This paper measures, in a controlled experiment, the point where all three meet: strong models take the developer-facing error text that wrappers inherit as an instruction, and stop.

Reviewer's judgment

First, we judge the most important number here to be not 69 points but the 82% under "cause alone." The agents could fix the problem all along. Failure came not from model capability but from one helpful-sounding sentence, which shows that the intuition "more guidance = better results" can be wrong in prompt and tool design.

Second, the fact that upgrading the model makes the same server text more costly matters operationally. An agent that works today could see its expired-credential recovery collapse from a model version upgrade alone, so model-swap regression tests must include error-path scenarios. An eval set that covers only success paths will not see this regression.

Third, the two remedies sit in different places and should be kept apart in practice. If you run the server, rewriting the text as a server tool call is the surest fix (84% on credentials, 88% on rate limits). If you connect other people's servers, a step-deletion filter is cheap (0.09 USD to process all 949 messages) and effective. But on missing permissions and rate limits the cause alone barely recovers (0% and 5%), so applying the filter across the board would delete correct steps too — the authors themselves recommend limiting it to credential errors.

Fourth, this connects to the semantic-layer debate. What an agent can do (its tool list) and what it is told to do on failure (error text) need to be aligned in the same vocabulary, so we see error text as part of the tool ontology that should be designed deliberately.

Practical levers

  • Name server tools in error text — instead of "please re-authorize," write "call auth_login first"; instead of "wait before retrying," write "wait a few seconds and call place_order again." It is the only form that also works for human developers.
  • Move human guidance to a separate field — put terminal commands, config paths, and web links in a field shown to human operators, and leave only the cause and tool calls in the text the model reads (the same idea as RFC 9457 problem details).
  • Filter third-party servers on credential errors only — insert a one-sentence prompt that deletes "next step" sentences only when a tool response is a credential error. Don't apply it to permission or rate-limit errors.
  • Add error paths to model-upgrade regression sets — keep expired-session, missing-permission, and rate-limit scenarios per failure type, and track the "ended without repair" and "handed back to user" rates.
  • Audit error text before connecting — before adding a new MCP server, check its source or sample responses for terminal, config, or web instructions. By this survey's count, 93% of credential-error steps (62 of 67) contained one.

Conclusion

This paper shows with public data that an MCP agent's failure can come from a single sentence the server returns rather than from the model, and that the cost can grow with newer models. The results are limited to OpenAI models and the explanation of model behavior is still a hypothesis, but the fixes are simple and the measured effects large, so both server and agent developers can apply them right away.

What operations teams should check as agents connect to more MCP servers and semantic layers continues in our blog post from the same day, Google's Gemini Agent Connects to Any MCP Server and a Knowledge Catalog.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…