Source Document

Mukul Chhabra, Shail Patel, Luigi Medrano, "CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production", arXiv:2609.30471 [cs.CL], submitted 2026-09-24, DOI 10.48550/arXiv.2609.30471. Affiliation: Dell Technologies.

This is a preprint without peer review (cs.CL). All three authors work for Dell Technologies, and the system under evaluation — a "production multi-agent assistant for enterprise technical support" — reads, in context, as the company's own live deployment. That is a clear conflict of interest: a company's own authors publishing an evaluation methodology for that same company's system. On the other hand, the paper does not promote that system; it discloses a structural flaw in the existing evaluation approach and an unresolved limitation in its own proposed fix (only 20% recall on procedural errors), which counts toward credibility.

Study Overview

Two questions drive the paper. First, why and how much does a literal reference-based LLM judge fail when evaluating production agents? Second, can that failure be fixed in a way that is verifiable without relying solely on expensive expert labels?

The data comes in three layers: a 30-item golden set from Dell's enterprise technical-support multi-agent deployment (23 entity-free policy answers, 7 entity-specific answers about cases, work orders, or assets); a stratified production sample of 400 traces from the same deployment, whose expert-labeling and gate-calibration analysis was preregistered but is not yet reported in this version; and a purpose-built diagnostic benchmark, Cargo-Bench, with 246 items generated from ten seeds.

The methodological core is a concept the authors call reference–instance divergence (RID). They decompose an answer into a procedure and instance-specific parameters, and show formally that a literal judge cannot separate the two: when entity values differ (a different case number, say), the judge penalizes the answer as wrong or hallucinated even when the underlying procedure is correct. Their fix, Cargo, reinterprets the reference as a procedural exemplar rather than a target answer, assigns each claim a three-way status — supported, contradicted, or unverifiable — against the observed live context, and abstains from scoring when retrieval confidence is low. To solve the problem that expert labeling is expensive and, more importantly, cannot tell a judge that is merely lenient from one that is actually better, the authors built a synthetic benchmark with ground truth fixed by construction and a discrimination index (DI = true-positive rate minus false-positive rate).

Key Results

Cargo-Bench's 246 items were scored by two judge models (gpt-oss-120b, mean of 3 seeds; gpt-oss-20b, 1 seed), producing 7,872 valid judgments — this figure covers only the 246-item main study (8 conditions × judges × seeds); a separately reported figure of 8,172 adds 300 judgments from a rubric-swap control, so the two totals cover different scopes and should not be conflated. Item families split into ones that should not be penalized (transplant P1, unverifiable P3) and ones that should (injected contradiction P2, procedural corruption P4); DI is the true-positive rate minus the false-positive rate.

Conditiongpt-oss-120b DI [95% CI]gpt-oss-20b DI [95% CI]
Direct (literal reference match).00 [.00, .00].00 [.00, .00]
Direct+Ctx (context added).00 [.00, .00]-.01 [-.03, .00]
NoRef (no reference).24 [.13, .35].39 [.27, .50]
Cargo.58 [.48, .68].56 [.46, .65]
Cargo+Auth (post-hoc fix).57 [.47, .67].58 [.48, .68]

Direct judging penalized all 50 transplant (P1) items and all 96 unverifiable (P3) items across both judges and every seed, giving a DI of essentially zero — the judge carries no information about actual quality. Direct+Ctx, which hands the judge the exact same live facts, failed identically on 120b and shifted only one item on 20b. Information alone does not fix this; what matters is telling the judge what that information is for.

Family (120b)TargetDirect penalty rateCargo penalty rate
P1 transplant (correct)should not penalize100% (50/50)0% (0/50)
P3 unverifiable (neutral)should not penalize100% (96/96)3% (3/96)
P2 injected contradictionshould penalize100% (50/50)100% (50/50)
P4 procedural corruptionshould penalize100% (50/50)20% (10/50)

Cargo cut false penalties to 0–3% while holding contradiction recall at 98–100% (50/50 on 120b, 49/50 on 20b). But it caught only 20% (10/50) of procedural corruptions on 120b, and the rate varied sharply by type: 10/26 step deletions (about 38%), 0/10 negations, 0/14 policy swaps. The rule "absent from context means don't penalize," intended for entity values, over-generalized to procedural claims. An explicit post-hoc fix targeting exactly this (Cargo+Auth) did not help (ΔDI=-.007, 95% CI [-.038, .023], straddling zero). A rubric-swap control on a 150-item subset found that swapping in the context-grounded scoring definitions alone — without the exemplar-framing preamble — already reached DI .49, showing most of the effect comes from redefining the rubric rather than from the framing.

The same blind spot survived human-style intervention. In an "LLM-as-annotator" study using written guidelines, a worked failure example, two independent passes, and adjudication, annotators recovered all 40 P1–P3 verdicts but caught only 3 of 10 procedural corruptions (30%). This is not an artifact of single-pass judging; it is structural.

Credibility Assessment

This is an unreviewed preprint, and as noted above there is a clear conflict of interest between the authors' employer and the system under evaluation. In credibility's favor: the authors openly disclosed their own proposal's shortcoming (only 20% procedural recall, unfixed by a post-hoc patch), and they used judge models (gpt-oss-120b/20b) distinct from the production assistant's own generator to limit self-preference effects.

The caveats run deep. The 246 Cargo-Bench items are synthetic — generated by one LLM, checked by a second, and spot-checked by the authors on only 40 of them. The deployment is a single company's single domain (enterprise technical support), so generalization to other industries or other judge models is not guaranteed. The paper's headline production analysis — expert agreement and gate calibration on 400 real traces (hypotheses H1, H2, H4) — was preregistered but is not yet reported in this version; every number available so far comes from the synthetic benchmark and the LLM-as-annotator study, not from human expert agreement. Only 7 of the 30 golden-set items are entity-bearing, which limits the statistical footing of the results that matter most.

Related Work

Compiled from the bibliographic details in this paper's own reference list. Session egress was blocked, so these three works could not be independently re-verified.

  • Gu et al. (2024). A Survey on LLM-as-a-Judge — prior work and overview. This survey catalogs known judge biases — position, self-preference, verbosity, task-dependent alignment. CARGO explicitly distinguishes itself: that catalog covers biases in the comparison mechanism, while RID is a mismatch in the comparison target itself.
  • Bavaresco et al. (2024). LLMs Instead of Human Judges? — prior work. A broad empirical study of whether LLM judges can replace human judges across 20 tasks. CARGO's contrast: such validation would miss a judge with near-zero discrimination unless it includes conditions where reference and instance diverge, as production agents routinely do.
  • Dubois et al. (2024). Length-Controlled AlpacaEval — extension and contrast. Addresses a different judge bias (verbosity) by correcting scoring rubrics. CARGO highlights the difference: that line of work fixes the comparison mechanism, while Cargo reinterprets the comparison target itself.

All three works address LLM-judge reliability, but none formalizes the idea that the reference answer itself can be the wrong target for a given instance — which is exactly the gap CARGO claims to fill. It is, in effect, a sixth failure mode alongside the five judge biases (position, verbosity, self-preference, format, drift) covered in LLM Judge Calibration.

Reviewer's Judgement

First, the most practically important result here is not DI .58 but the failure of Direct+Ctx. Handing the judge live facts is the first fix most teams would try, and it did nothing on 120b. Adding context to a prompt is not the same as telling the judge what that context is for, and teams should not assume that adding facts alone solves a structural mismatch.

Second, reading the 20% procedural-recall figure as simply "Cargo's failure" is too quick. Because the authors tried a post-hoc fix and reported transparently that it did not work, this limitation is a known, shippable trade-off rather than a hidden defect. Having a single number — the discrimination index — is precisely what let them tell "more lenient" apart from "actually better." That measurement tool may outlast the specific Cargo implementation.

Third, the fact that the LLM-as-annotator study — with written guidelines, two independent passes, and adjudication, mimicking a human review process — reproduced the same blind spot (3/10) supports the authors' own conclusion: this is not a prompt-engineering problem but one that needs a structurally different check, such as per-step verification against the reference. Attempts to fix this purely by rewording prompts are likely to keep hitting the same wall.

Putting It to Work

  • Reframe the reference as a procedural exemplar — remove the implicit "different value means wrong" assumption from judge prompts, and state explicitly that the reference shows only procedure, policy, and structure.
  • Adopt three-way claim status — stop treating values absent from observed context as automatic hallucinations; split into supported, contradicted, and unverifiable, and penalize only contradictions.
  • Gate evaluation on retrieval confidence — abstain from scoring traces with no appropriate reference, and treat that gap as a signal for human review or golden-set expansion.
  • Cover procedural errors with a separate check — assume a Cargo-style judge can miss procedural corruptions, and run a step-by-step comparison or a no-reference (NoRef) score in parallel on the procedural dimension.
  • Validate the judge itself with a discrimination metric — when adopting a new scoring method, track true positives alongside false positives so you can tell "more lenient" from "actually better."

Conclusion

This paper's contribution is not a new scoring model but the formalization and measurement of how the LLM judge's implicit assumption — that the reference answer is the target — structurally breaks down in production agents where every instance has different entities. Cargo lifts the discrimination index from 0 to .58, but procedural-error detection remains stuck at 20%, and a post-hoc fix did not close that gap. Given the conflict of interest between the authors' employer and the evaluated system, and given that the production expert-labeling results are not yet published, these numbers are best read as a directional signal, with procedural errors covered by a separate check. For a different angle on judge bias, see LLM Judge Calibration: Five Bias Types and How to Correct Them.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…