From 5.7% to 1.9%: What Contextual Retrieval Actually Proved
Anthropic's Contextual Retrieval experiment, published in September 2024, measured a 5.7% failure rate for top-20 chunk recall using standard embedding search. Prepending 50–100 tokens of chunk-specific context before embedding — Contextual Embeddings alone — cut that to 3.7%, a 35% reduction. Building the same contextualized text into a BM25 lexical index alongside the embeddings brought it to 2.9%, a 49% reduction. Adding a final stage — reranking the top 150 candidates with Cohere's reranker down to the 20 passed to the model — took it to 1.9%, a 67% reduction. Generating that context costs about $1.02 per million document tokens at 800-token chunks and 8k-token documents, and prompt caching cuts the latency by more than half and the cost by up to 90%.
The New Boundary Where Search Becomes a Tool Call
The MCP spec treats resources as data exposed "without significant computation or side effects," while tools are model-controlled calls the model decides to invoke and how to parameterize. Wrap a RAG pipeline in a single MCP tool, and the host model stops treating the response as ordinary text — it treats it as a verified tool result. An earlier post covered the schema and call-limit side of that wrapping; whether the 1.9% recall-failure number still holds for this specific call is a different question entirely — if the index is mid-rebuild or reranking times out, the model accepts the degraded result with the same confidence.
From Design to Operations: A Trust Checklist for RAG Exposed via MCP
(Planning) Set targets around the tool response's freshness and scope, not the retrieval algorithm itself. Cap index-rebuild lag at p95 under 10 minutes, hold the tool response's recall-failure rate at 2% or less — using Anthropic's reranked 1.9% as the baseline — and declare 0% exposure for chunks outside a user's permission scope. Without these three numbers fixed in code, a tool ships on the impression that "search looks fine," with no canary to catch otherwise.
(Failure patterns) Three show up repeatedly: index lag, where a document update hasn't propagated to the vector index yet and the tool returns a stale chunk; permission-scope leakage, where a chunk outside the user's access scores high enough on reranking to land in the top 20 and ships in the tool response anyway; and silent degradation, where a reranker call times out and the server falls back to passing the pre-rerank 150 candidates through unchanged, quietly pushing the failure rate back toward 5.7%.
(Recovery branches) For index lag, attach a document-version timestamp to the tool response's metadata so the host model knows how stale the index is, and invalidate the cache and re-query once that lag crosses a threshold. Block permission-scope leakage by applying the user's access filter before reranking, not after, so the candidate pool itself is already narrowed. Treat a reranker timeout as a failure rather than disguising it as success — fall back to a safe reduction that asks for user confirmation instead of silently returning the unranked 150.
(Operations checklist) Before deployment, re-measure recall@20 failure against a golden query set and compare it to the 1.9% baseline. Standard log fields — query hash, returned chunk IDs, reranker score, index version, latency — let you tell immediately, when a quality complaint comes in, whether the index or the reranker is at fault. Mask or replace PII such as customer names and emails inside chunks before they leave through the tool response.
(Improvement loop) Collect recall failures weekly to see which query types hit the baseline most often, and keep index-side changes — embedding or reranker swaps — in a separate changelog from tool-schema changes. When quality drops, cross-referencing the two logs tells you immediately whether an embedding refresh or a tool-definition change caused it.
Takeaways to Apply Now
Exposing RAG as an MCP tool means holding the line at a 2% recall-failure rate, a 10-minute p95 index-rebuild lag, and 0% permission-scope leakage. Design index lag, scope leakage, and reranker timeouts as three separate recovery branches, and no amount of retrieval-layer sophistication will quietly leak away at the tool-call boundary.
References
Introducing Contextual Retrieval — Anthropic
Resources — Model Context Protocol Specification
Ask AI about this article
The assistant has read this article. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…