Only Six in Ten Finish: Deployment Criteria for Long-Horizon Browser Agents
Even the top open-source browser agent finishes only 65.1% of realistic long-horizon tasks — deployment calls need step-length segments, not one average.
- On BrowserBench (350 tasks, avg 37.9 steps), the top open-source 27B model hits 65.1% and even GPT-5.5 tops out at 67.4% — a 32.6pp miss rate
- Online RL (DAO-GRPO) gains scale with difficulty: +12.4pp on hard tasks, +13.4pp (13.3%→26.7%) on tasks over 50 steps, only +1.0pp on easy ones
- Mixing recovery data too early in the curriculum drops success to 35.1% vs. 38.0% with proper ordering — budget overruns should count as failures













































































































































