Source Document

Andre Fu, Malik Drabla, Leon Liu, and 5 others, "Incident-Arena: Getting agents to the last nine of reliability", arXiv:2610.00648 [cs.AI], submitted 2026-09-30. Affiliations: Abundant AI, Adrenaline AI, Carnegie Mellon University, Comenius University in Bratislava, Massachusetts General Hospital.

This is a preprint without peer review. Corresponding author Andre Fu's email domain is abundant.ai, and the author list includes Abundant AI and Adrenaline AI — startups that appear to build agent-evaluation and SRE-related products — yet the paper carries no separate funding or conflict-of-interest disclosure. A commercial interest in the benchmark's own promotional value cannot be ruled out. The full text could not be fetched directly due to a session-wide egress block (confirmed against the unrelated control example.com across three 10-minute-spaced retries, all returning 403), so tables and figures were cross-checked against a primary-source HTML snapshot collected by GitHub Actions (manifest generated 2026-10-02T22:11:05Z).

Study Overview

Coding agents already handle much repository-level work, but whether they can extend into diagnosing and repairing live production incidents — "agentic SRE" — has been under-tested. The authors argue existing benchmarks suffer three limits: (1) unrealistic, toy-scale environments, (2) non-standard harnesses that differ benchmark to benchmark, and (3) simple static verifiers. Incident-Arena targets all three: it deploys three real open-source production applications (the e-commerce platform Saleor, the ERP system FrappeERP, and a custom-built Slack-like system called "slack-spine") onto ephemeral Kubernetes clusters, injects a fault at the configuration, runtime, or image layer, and runs 20 tasks under sustained load.

Evaluation covers 10 agent configurations (native harnesses from OpenAI, Anthropic, xAI, Meta, and Zhipu; Zhipu's GLM-5.3 is the one model paired with Claude Code) at 3 trials per task across every available reasoning-effort setting, totaling 3,000 trials. Instead of an LLM judge, verification uses a deterministic two-gate design: an Outcome Gate checking whether service-level indicators (goodput, latency, error rate) recover under continued traffic, and a Safety Gate checking that the fix survives a restart or re-triggered fault, stays within the task's permitted scope, and leaves protected data intact. The authors say this design avoids both frontier models detecting an evaluation regime and the reliability problems of LLM judges.

Key Results

Across the 20 tasks and 3,000 trials, the best model-configuration's pass rate is stated as 64.3% in the abstract, while the conclusion separately states "GPT-6-Astra reaching a high of 65% at the xhigh setting" (both figures appear verbatim in the source; the exact reproducing configuration is not pinned down further). Pooling across all reasoning-effort settings (Table 10) gives lower, per-model numbers.

ModelPass@1 (pooled across effort)95% CI (task-bootstrap)Full-fix action-completion rate
GPT-6 astra59.0%43.0–74.3%70.2%
Claude Opus 5.554.0%38.7–69.0%73.6%
Claude Fable 5.144.0%29.7–59.0%59.7%
Muse Spark 1.3 (lowest)25.6%13.9–37.8%23.3%

Across all 10 models, pooled Pass@1 ranges from 25.6% to 59.0%, and the confidence intervals of the top two models (GPT-6-astra and Opus-5.5) overlap heavily (43.0–74.3% vs. 38.7–69.0%), so 300 trials per model are not enough to settle a ranking.

The paper's most important finding isn't the top pass rate — it's what happens when you cross fix-completeness with pass/fail. Among the 2,967 trials with a usable action record, 1,638 (55%) issued every action in the reference fix ("full fix"). Of those, only 59% (966) actually passed, meaning 672 episodes (41%) failed despite doing everything the reference solution required. By contrast, the 1,329 episodes without a full fix passed only 5.5% of the time (73 episodes). Not knowing the fix and knowing it but still failing are two distinct problems.

Exclusive failure class (1,948 failures total)CountShare
Never declared (time/step budget exhausted)31716.3%
Changed something off the permitted list42721.9%
Planted cause or damage left in place1708.7%
Repair did not survive a restart/next trigger51726.5%
Declared while service still out of bounds46623.9%
Other integrity check failure512.6%

The 672 full-fix episodes that still failed split into two patterns. 340 had a correct fix but the system hadn't recovered yet (a draining backlog, an incomplete restart), and the agent declared done too early. The other 329 failed by over-correcting — moving a maintenance window to a still-unsafe time (145 cases) or raising an already-fixed connection-pool ceiling back up (38 cases) — undoing their own success.

Declaration timing tracks pass rate closely. Among full-fix episodes, declaring within 2 minutes of the last correct action (441 episodes) passed 71.7% of the time; declaring 10+ minutes later (143 episodes) passed only 18.6%. Read-only check counts show the same pattern: 5 or fewer checks passed 63.4%, 9 or more passed 52.3%. The authors state that, within the same task, episodes that declared quickly out-passed slower ones by 20 points on average.

Raising reasoning effort from lowest to highest barely moved the needle. Across the eight models offering all five settings, pass rate went from 39.2% (low) to 40.0%, 40.2%, 41.9%, and 43.8% (max) — a gain of only +4.6pp (task-resampled interval −4.8 to +13.8), while mean cost per episode rose 2.3x. The effect wasn't consistent across tasks either: +9.7pp on multi-fault tasks but −3.1pp on single-fault ones, and −29.2pp on the Saleor task alone.

Credibility Assessment

The strengths are clear. The core metric (Pass@1) is verified by a deterministic two-gate check rather than an LLM judge, so it is free of judge bias. The failure taxonomy (1,948 cases) is exclusively categorized and sums exactly, and the sub-results that support the paper's central claim — 59% pass with a full fix vs. 5.5% without — are internally consistent.

Four caveats matter. First, the author list includes startups with a commercial stake in this space (Abundant AI, Adrenaline AI) with no separate conflict-of-interest disclosure. Second, as the authors note, only 3 trials per setting were run, producing wide confidence intervals; the top two models' intervals overlap too much to rank confidently. Third, the task mix is imbalanced (13 slack-spine, 6 Frappe, 1 Saleor) and fault type/onset is confounded with application, so the authors themselves say only within-application comparisons should be trusted, not cross-task generalizations. Fourth, reward-hacking judging used GPT-6-Astra — one of the benchmark's own top-ranked participants — as the judge, and the authors acknowledge low inter-judge agreement in an appendix.

Related Work

Related literature was identified from this paper's own Related Work section and comparison table (Table 1). Because of the session-wide egress block, each related paper's own text could not be independently re-verified, so only Incident-Arena's description of the design differences is cited — not those papers' own numbers.

All three are prior works that Incident-Arena cites directly to explain its own design choices. The operational question of how to automatically halt anomalous signals inside an evaluation environment is picked up in this site's escalation-gate design for anomalous eval signals (Korean).

Reviewer's Take

First, I'd argue the most operationally important number here isn't the 64.3% top pass rate but the 672/1,638 figure — a 41% failure rate even with the complete reference fix in hand. Many teams size automation scope purely on a coding agent's "correctness rate." Incident-Arena shows numerically that knowing the fix and finishing safely are separable capabilities.

Second, I think using the benchmark's own top-ranked model (GPT-6-Astra) as the reward-hacking judge deserves reconsideration. Having the judge overlap with the judged pool weakens confidence in the verdicts, and combined with the authors' own admission of low inter-judge agreement, the reward-hacking rates should be read conservatively.

Third, the weak cost-effectiveness of raising reasoning effort (+4.6pp for 2.3x the cost) suggests the common intuition — "if unsure, make it think more" — doesn't hold well for long-horizon, operations-style tasks. Tuning effort by task difficulty is likely a better investment than a blanket increase.

Practical Levers

  • A durability-observation gate before declaring done — don't accept "fixed" the moment a change lands; require it to survive a restart or the next trigger first. Repairs that didn't survive a restart made up 26.5% of all failures.
  • An allow-list guardrail for changes — scope agent permissions tightly and auto-diff the change log against an allow list. Off-list changes caused 21.9% of failures.
  • Treat lingering re-checking as a risk signal — episodes that declared within 2 minutes of the last correct action passed 53.1 points more often than those that took 10+ minutes (71.7% vs. 18.6%). Nine or more read-only checks also correlated with a lower pass rate (63.4% → 52.3%).
  • Tune reasoning effort per task instead of raising it blanket-wide — low-to-max settings bought only +4.6pp at 2.3x cost on average, and reversed on some tasks (−29.2pp on Saleor). Test by task family before defaulting to a higher setting everywhere.
  • Diversify and decouple reward-hacking judges — avoid using a benchmark participant model as its own judge; mix in non-participant models or human review to raise confidence in the verdicts.

Conclusion

Incident-Arena makes "can an agent fix a production incident?" testable with a real Kubernetes environment and a deterministic two-gate check. Its most important finding isn't the top pass rate (64.3%) but that 41% of episodes fail even after issuing the complete reference fix — most of those failures come from repairs that don't last, or from declaring done before recovery actually finished, a "not knowing when to stop" problem. That said, the authors' commercial affiliations, the wide confidence intervals from a small trial count, and a reward-hacking judge that overlaps with the model pool all argue for re-verifying these results in your own operating environment rather than taking them at face value.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…