Source
Xing Han Lù, Dheeraj Vattikonda, Sina Hajimiri, Fatemeh Pesaran Zadeh, Parishad BehnamGhader, Ghazwa Darwiche, Amirhossein Kazemnejad, Christopher Pal, Alexandre Drouin, Siva Reddy, "AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks", arXiv:2610.11050 [cs.AI], submitted 2026-10-08, DOI 10.48550/arXiv.2610.11050. Affiliations: ServiceNow Research, McGill University, Mila – Quebec AI Institute, ÉTS Montréal, Seoul National University, Université Laval, Polytechnique Montréal, Canada CIFAR AI Chair. Benchmark and results fully released (github.com/ServiceNow/agenthorizon).
This is an arXiv preprint that has not gone through peer review at a conference or journal. Funding is a mix of public and corporate sources: Canada's NSERC, the Canada CIFAR AI Chair program, the ServiceNow–Mitacs Accelerate program, and an RBC Borealis AI fellowship. Several authors are affiliated with ServiceNow Research, which is a real conflict-of-interest disclosure point. That said, every judge model evaluated — GPT, Gemini, Claude, Qwen, Gemma — is a third-party model; no ServiceNow agent or judge is benchmarked, so the direct incentive to flatter a house product is weak. Session egress was blocked, so the full text was not read via live WebFetch; verification instead used a GitHub Actions snapshot of the arXiv HTML full text (2610.11050v1) captured at 2026-10-09T22:11:46Z.
Study Overview
Two questions drive the paper: (1) can LLM judges catch computer-use trajectories that look complete but actually violate the instruction or cause an unwanted side effect, and (2) does giving a judge tools (code execution, selective screenshot inspection) — "agentic judging" — beat showing it the whole trajectory at once ("direct judging")?
The authors recorded 166 hours of human demonstrations across three operating systems, then built paired instructions that are similar but incompatible — pairing a recording with the swapped instruction turns it into a failure case by construction. After quality and evidence review, this yielded 1,373 instruction-trajectory pairs: 523 positives and 850 adversarial negatives. These were split by difficulty into an evaluation set AgentHorizon (AH, 528 items), a simpler complement AgentHorizon-Simple (AH-S, 683 items), and a development set AgentHorizon-Development (AH-D, 162 items) — the positive-plus-negative counts for each split match the stated totals exactly on verification. Eleven judge models (six closed, five open-weight) were evaluated two ways: as agentic judges across five harnesses (Claude Code, Codex, Gemini CLI, OpenHands, OpenCode) that inspect the trajectory with tool calls, and as direct judges given a fixed screenshot representation in a single chat-completion call.
Key Results
Accuracy is reported as balanced accuracy — the mean of positive-class and negative-class accuracy, so a judge that always predicts one label scores 50%. On the harder AH split, GPT-5.5 with the Codex harness scored highest at 80.9%, followed by Gemini 3.1 Pro with Gemini CLI (77.7%) and Claude Opus 4.7 with Claude Code (76.0%). The best open-weight judge was Qwen3.6 27B with OpenCode at 70.4%. But the same model's score swings sharply depending on how it accesses the trajectory.
| Model | Interface | AH Bal. | AH Pos | AH Neg | AH-S Bal. |
|---|---|---|---|---|---|
| GPT-5.5 | Codex (primary) | 80.9 | 71.8 | 90.0 | 92.6 |
| GPT-5.5 | OpenCode | 77.2 | 70.9 | 83.4 | 92.2 |
| Qwen3.6 27B | Codex | 71.6 | 70.5 | 72.8 | 94.5 |
| Qwen3.6 27B | OpenCode (primary) | 70.4 | 75.3 | 65.4 | 93.6 |
| Qwen3.6 27B | OpenHands | 51.2 | 58.1 | 44.2 | 89.0 |
GPT-5.5 gained 12.3 points moving from direct judgment (68.6%) to the Codex harness (80.9%), while Qwen3.6 27B lost 7.2 points moving from direct judgment (77.6%) to the OpenCode harness (70.4%) — the assumption that "giving a judge tools helps" cuts in opposite directions depending on the model. Comparing the same model across three harnesses (Qwen3.6 27B, above) shows up to a 20.4-point swing from harness choice alone.
A more concerning number than balanced accuracy is mistake-type recall, measured over the full 1,373-item benchmark. "Critical Mistake" (the primary goal was not achieved) and "Misunderstanding" (a plausible but incorrect, harmless interpretation) are caught at roughly 60-90% rates. But recall for "Bad Side Effect" — the goal was achieved, but the trajectory caused harm, risk, or an unintended consequence — tops out at just 24.1% for the best judge (GPT-5.5+Codex), and two open-weight models recorded exactly 0.0%.
| Model | Interface | Critical Mistake recall | Bad Side Effect recall | Misunderstanding recall |
|---|---|---|---|---|
| GPT-5.5 | Codex | 77.7 | 24.1 | 72.9 |
| Claude Opus 4.7 | Claude Code | 75.8 | 12.1 | 71.4 |
| Gemini 3.1 Pro | Gemini CLI | 57.5 | 7.3 | 87.3 |
| Qwen3.6 27B | OpenCode | 60.4 | 5.6 | 74.3 |
| Gemma4 26B-A4B | OpenCode | 47.6 | 0.0 | 18.6 |
In other words, the 80.9% balanced-accuracy headline tells you almost nothing about a judge's ability to catch "succeeded but caused harm." The authors also isolated which inputs actually matter via ablation: dropping screenshots and leaving only the text action log cuts balanced accuracy to 57.7%, and removing the instruction itself (making the task provably unsolvable by construction) drops it to near-random 47.8-49.8%. Text logs alone carry too little signal; screenshots are where most of the judging evidence lives. On cost, GPT-5.5+Codex averages $0.40 per trajectory and Claude Opus 4.7+Claude Code averages $0.66, while self-hosted Qwen3.6 27B+OpenCode reached 70.4% balanced accuracy with no marginal API cost (and under half the input tokens of most closed models — 243K on average).
Credibility Assessment
Three things support trust here. The paired design isolates judge error from task-difficulty confounds by swapping only the instruction. The final 523 positives survived quality and evidence review out of 1,700 initial candidates. And the balanced-accuracy arithmetic checks out on every row when recomputed from the reported class accuracies — e.g., GPT-5.5+Codex: (71.8+90.0)/2 = 80.9; Qwen3.6 27B+OpenCode: (75.3+65.4)/2 = 70.35 ≈ 70.4.
Caveats are real too. This is an unreviewed preprint with the ServiceNow funding and affiliation conflict noted above. The authors' own stated limitations: no independent human inter-annotator agreement was collected (only consistency across review stages), the corpus is English-only across three operating systems using screenshots and action logs rather than full application state, and the paired design targets subtle near-misses rather than general agent failures like loops or gradual drift. As contrasting evidence, the same lead author's earlier benchmark AgentRewardBench (below) reported that even the best judges on more general web-agent trajectories stayed under 70% precision versus human experts — a different metric (precision vs. balanced accuracy) that can't be compared numerically, but both studies converge on the same conclusion: LLM judges are not yet reliable enough for unsupervised automation.
Related Work
- Lù et al. (2025). AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories — the same lead author's earlier benchmark: 1,302 web-agent trajectories across 5 benchmarks and 4 LLMs, evaluated by 12 LLM judges, found the best judge still under 70% precision versus human experts. This paper extends the same concern to long-horizon, multi-application computer-use tasks and adds a paired adversarial design that surfaces the "side effect" failure type (prior work, extension).
- Zhuge et al. (2024). Agent-as-a-Judge: Evaluate Agents with Agents — introduced the "agentic judging" methodology this paper's tool-using judges build on, reporting that agentic judging substantially outperformed plain LLM judging on the 55-task DevAI benchmark. This paper's long-horizon results partly contradict a blanket claim of superiority, since tool access improves some models and hurts others (prior work, partial contradiction).
- Kuntz et al. (2025). OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents — measures computer-use agent safety (misuse, prompt injection, misbehavior) with an automated LLM judge that agreed with human annotators at 0.76-0.79 F1. This paper targets success/failure judgment with a harder paired-adversarial design, under which side-effect detection recall falls to as low as 24.1% — a more conservative conclusion on the same underlying question of judge reliability (related work, methodological contrast).
Taken together, the reliability problem for LLM judges shows up repeatedly across web agents (AgentRewardBench), safety judging (OS-Harm), and long-horizon computer use (this paper) — and the stricter the evaluation design (paired adversarial pairs), the more clearly a judge's weaknesses surface.
Reviewer's Assessment
First, the number that matters most here is not 80.9% but the 24.1% Bad Side Effect recall. Balanced accuracy weights positives and negatives equally, but in practice the hard-to-reverse harms — a payment to the wrong account, an action on the wrong device — concentrate in the Bad Side Effect category. A ceiling of roughly 24% recall in that category means that using today's LLM judges as an unsupervised approval gate risks letting "successful but harmful" trajectories straight through.
Second, the fact that harness effects run in opposite directions by model (helping GPT-5.5, hurting Qwen3.6 27B) undercuts any blanket claim that agentic judging beats direct judging. The authors themselves note this comparison changes the whole harness at once (screenshot selection, context management, tool exposure) without isolating individual factors — which argues for validating the exact model-plus-harness combination you plan to deploy, not the model alone.
Third, an open-weight model (Qwen3.6 27B) beating most closed models at 77.6% in direct-judging mode with zero marginal API cost is a reasonable argument that the most expensive judge is not automatically the best choice. But since the ablation shows open-weight scores are sensitive to screenshot preprocessing, that decision should weigh preprocessing choices too, not cost alone.
Practical Takeaways
- Validate the model-plus-harness pair, not the model alone — the same model's balanced accuracy can swing more than 20 points depending on tool access, so test the exact combination you'll deploy.
- Track side-effect detection as its own metric — balanced accuracy alone hides Bad Side Effect recall rates in the single digits to low teens. Pipelines with hard-to-reverse actions should monitor this recall separately.
- Keep visual evidence (screenshots) in the judging pipeline — text action logs alone drop balanced accuracy to 57.7%.
- Consider self-hosted open-weight judges as a cost alternative — Qwen3.6 27B matched or beat most closed judges (77.6% direct, 70.4% via OpenCode) at zero marginal cost.
- Report positive- and negative-class accuracy separately — balanced accuracy alone can't distinguish over-acceptance (false positives) from over-rejection (false negatives).
Conclusion
This paper's contribution isn't a new judge model — it's quantitative evidence that current LLM judges fall short of a trust bar for long-horizon computer-use trajectories. Even the best result (GPT-5.5+Codex at 80.9%) is far from perfect, and more importantly, recall for the practically riskiest failure type — success with a harmful side effect — tops out around 24%. That tool access helps or hurts depending on the model also undermines the assumption that agentic judging is strictly better. Before treating automated judging as the final QA gate for deployed agents, validate the exact model-plus-harness combination on your own tasks and track side-effect detection separately. A companion piece on how aggregate scores can mask real failure patterns in agent evaluation design is available at An Agent Eval Set Design Checklist.
References
- AgentHorizon — arXiv abstract (primary source)
- Same paper, HTML full text — used for table/number verification
- AgentRewardBench (Lù et al., 2025) — related work source
- Agent-as-a-Judge (Zhuge et al., 2024) — related work source
- OS-Harm (Kuntz et al., 2025) — related work source
- An Agent Eval Set Design Checklist — sunny34.com blog
Ask AI about this review
The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…