Source Document

Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu, Lucy Lu Wang, "HERA: Harness–Environment Co-Evolution for Reliable Agentic Abstention", arXiv:2610.06563 [cs.AI], submitted 2026-10-05. Affiliation: University of Washington, University of Leeds, Carnegie Mellon University (Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang); Stanford University, Allen Institute for AI (Pan Lu, Lucy Lu Wang). Project page released (hera-bench.github.io).

This is a preprint without peer review. The Ethics Statement cites the "ICLR Code of Ethics," suggesting an ICLR submission, but the paper does not state an acceptance outcome or venue, so it should be read as a draft. All authors are affiliated with universities or nonprofit labs; no overlap with the companies behind the evaluated models, and no funding disclosure, was found in the text. Session egress was blocked, so the full text was checked against the HTML snapshot fetch_source_snapshot.py collected at 2026-10-07T22:10:39Z.

Study Overview

The question is: measure not just whether an agent can complete a task, but whether it knows when to stop. The authors build a pipeline that auto-generates matched solvable/infeasible task pairs by applying controlled mutations to an environment — one variant stays solvable (Act), the other is missing information, permission, or evidence and should trigger a stop (Abstain). On top of this they run HERA, a harness–environment co-evolution loop: failures are analyzed to patch the harness (prompts, memory, tools, control logic) while also generating new, harder tasks that target those same failures for the next round. After six rounds, the training pool grows from 20 to 50 task pairs, and the final harness is checked on a separate held-out set, hera-bench (60 pairs, disjoint domains). The base model is GPT-5.6-Luna; the same final harness is then applied, without further tuning, to 19 other models.

Key Results

On held-out hera-bench, holding the base model (GPT-5.6-Luna) fixed and swapping only the harness: Abstain is the share of infeasible tasks correctly refused, Act is the share of feasible tasks completed, and Pair requires both to hold for the same matched pair.

MethodAbstainActPair
Base harness (H₀)61.7%68.3%48.3%
H₀ + abstention prompt68.3%63.3%46.7%
Harness-only evolution76.7%71.7%60.0%
Task-augmented evolution80.0%75.0%65.0%
Task gen. from H₀ failures78.3%75.0%73.3%
Co-evolution (HERA)83.3%76.7%70.0%

Simply adding a system-prompt instruction raises Abstain but drops Act from 68.3% to 63.3% — it trades away completions. HERA scores highest on both Abstain and Act among all compared methods, meaning the gain in stopping power didn't come from refusing everything. Note this table is the held-out set; on the cumulative training pool (T₆, 50 pairs) HERA reaches Abstain 86.0% / Act 88.0% — higher numbers that should not be quoted interchangeably with the held-out figures.

Applying the harness built on GPT-5.6-Luna to 19 other models lifts Abstain by an average of +15.3pp (range +6.7 to +40.0pp) and Pair by +12.5pp (+1.7 to +28.3pp). Even the already-strong GPT-6-Astra improves (Abstain 81.7%→91.7%, Pair 70.0%→80.0%). Weaker models gain less: Mistral Large 3 moves only from Abstain 6.7%→16.7% and Pair 3.3%→5.0%, while its over-abstention rate (wrongly refusing a solvable task) rises from 13.3% to 15.0%. On cost, across 14 models median API cost rises 2.09x and tokens 3.31x, yet GPT-5.6-Luna+HERA matches the base-harness GPT-6-Astra's 70.0% paired success at roughly $0.0091 per episode versus $0.0610 — about 85% lower.

Credibility Assessment

Three things support trust: the base model is held fixed while only the harness varies, held-out and training-pool results are reported separately rather than conflated, and the harness is cross-checked on 19 models rather than one. Caveats remain — the paper does not state how Abstain/Act/Pair judgments are scored (automated grader reliability against human judgment isn't reported), cross-model gains vary widely (+1.7pp to +40.0pp), and weaker open-weight models can see over-abstention worsen. All authors are university or nonprofit-lab affiliated, which lowers one obvious conflict-of-interest concern, but no funding statement was found. Contradicting evidence: Wang et al. (2026), cited below, questions whether harness-evolution gains generalize at all.

Related Work

Taken together, these three papers suggest that roughly half of paired agent-abstention tasks fail without intervention, and that whether harness-level fixes generalize remains contested — HERA's cross-model evidence is a data point in that debate, not a resolution of it.

Reviewer's Judgement

First, I'd argue the most practically important result isn't 83.3% but the second row of Table 1: adding an abstention prompt alone drops Act from 68.3% to 63.3%. That's the cheapest fix many teams are likely already using, and this shows it isn't free.

Second, Mistral Large 3's rising over-abstention (13.3%→15.0%) is easy to miss behind the "+15.3pp average" headline. Averages get pulled up by strong models' large gains, while weak models may see small gains paired with a growing refusal problem. Knowing where your own model sits on that cross-model table matters more than the average.

Putting It to Work

  • Measure abstention separately from completion — use matched feasible/infeasible pairs and report Abstain, Act, and Pair rather than a single success rate.
  • Treat prompt-only fixes as having a cost — a bare "stop if unsure" instruction can raise refusals while lowering completions; compare it against harness-level changes.
  • Pilot before porting a harness across models — run a small Pair/over-abstention check on your own model before adopting a harness tuned on a different one.
  • Compare cost per matched outcome, not raw token growth — the real signal is cost per successful episode at equal performance, not the multiplier in tokens used.
  • Watch over-abstention on weaker models specifically — track whether a harness meant to improve stopping also starts refusing solvable tasks.

Conclusion

HERA's contribution is shifting harness evolution from "adapt to a fixed task set" to "let the environment and harness evolve together." The headline numbers — 61.7%→83.3% abstention on held-out tasks, +15.3pp average across 19 models — point in a clear direction, but unresolved grader-reliability questions and a documented over-abstention regression on weaker models mean the numbers are a direction, not a specification. Pilot on your own model and tasks before adopting. For the operational side of gating a harness rollout on both completion gains and safety regressions, see Promotion Gates for a Self-Evolving Agent Harness.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…