Source Document

Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasović, "The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge", arXiv:2609.30604 [cs.CL], submitted 2026-09-24, DOI 10.48550/arXiv.2609.30604. Affiliation: University of Utah.

This is a preprint without peer review (cs.CL, no conference/journal acceptance noted). All authors are affiliated with the University of Utah and have no direct relationship with the vendors whose models were evaluated. Funding sources are worth flagging, though: Google funded the construction of the benchmark itself, and the artifact platform is Google Workspace (Docs, Sheets, Slides). The authors say they excluded Google's own models from evaluation specifically to mitigate this structural advantage. Anthropic and OpenAI also donated API credits for benchmarking — and their models (Claude Opus 4.7, GPT-5.5) are among those evaluated, so this is not entirely unrelated funding. Session egress was blocked when this review was produced, so the full text was cross-checked against a GitHub Actions snapshot of the primary-source HTML fetched at 2026-09-28T22:32:00Z (one day prior to this review).

Study Overview

One question drives the paper: can an agent fully play the role of an "assistant" — from finding information on the web to synthesizing, organizing, and presenting it as a finished artifact (a document, spreadsheet, or slide deck)? Existing computer-use benchmarks either test short, easily verifiable web-navigation tasks or ask agents to produce text-only research reports, and neither measures the last mile of real work: visual and spatial layout, formatting, and structure.

The authors built a new benchmark, KNOWS (Knowledge Navigation and Organized Web Synthesis). Its 110 tasks open with realistic natural-language instructions (204.4 words on average) and require an agent to use a live web browser, a vision-language model (VLM), and Google Workspace to produce an artifact — 25 documents, 40 slide decks, or 45 spreadsheets. Grading is hybrid: deterministic checks (structure, formatting, coordinates) combined with LLM/VLM judgment, broken into an average of 5.9 checkpoints and 24.6 evaluation steps per task (2,716 steps total). The evaluators' own reliability was validated against expert judgment: Cohen's kappa of 0.64 (82% agreement) overall, and kappa 0.69 (89% agreement) on the subset of steps decided purely by LLM/VLM judgment.

Key Results

Across seven model–harness combinations (three text-only, two multimodal, two AI browsers), only the best pairing — Perplexity Comet with Claude Opus 4.7 — achieved a nonzero full success rate (SR): 2.7% overall (12.0% on Docs). Every other combination scored 0%.

Model / HarnessSR (Overall)ASC (Overall)ACF (Overall)SF (Overall)
GPT-5.5 (text-only)0%20.6%41.9%38.0%
Claude Opus 4.7 (text-only)0%6.4%17.0%13.6%
DeepSeek V4 Pro (text-only)0%12.6%31.1%26.4%
GPT-5.5 (multimodal)0%21.8%49.5%46.7%
Claude Opus 4.7 (multimodal)0%16.7%42.5%39.1%
ChatGPT Atlas (AI browser)0%21.6%45.4%40.6%
Perplexity Comet (AI browser)2.7%35.4%70.0%64.7%

The metrics are not interchangeable. SR awards a point only for a fully completed task — a strict, all-or-nothing measure — while ASC, ACF, and SF give partial credit at the checkpoint or step level. Comet's ACF 70.0% / SF 64.7% look respectable in isolation, but the paper's own showcased example (Figure 3) scored ACF 0.54 / SF 0.59 while completing only 0.14 of its checkpoints (ASC) — the authors themselves describe the resulting artifact as "a visual mess."

Visual input mattered. Adding screenshots to Claude Opus 4.7 (multimodal vs. text-only) raised ASC by +10.3pp, ACF by +25.5pp, and SF by +25.5pp (checked: 16.7-6.4=10.3, 42.5-17.0=25.5, 39.1-13.6=25.5 — table and text agree). The harness effect was larger still: swapping only the harness for the same Claude Opus 4.7 (BrowserGym → Comet) raised SR by +2.7pp, ASC by +18.7pp, ACF by +27.5pp, and SF by +25.6pp (checked: 35.4-16.7=18.7, 70.0-42.5=27.5, 64.7-39.1=25.6 — consistent). Harness infrastructure — browsing stack, action space, tool interfaces — moved the needle more than the underlying model did.

Credibility Assessment

Three things support the results. First, the evaluators' own reliability was independently validated against expert judgment (kappa 0.64–0.69). Second, the study holds the model fixed while varying harness and input modality, partially separating model effects from harness effects. Third, aware of the Google Workspace training bias, the authors excluded Google's own models and ran a small cross-platform check on Microsoft 365 (5 tasks).

The caveats are just as real. As noted above, Google funded the benchmark's construction and Anthropic/OpenAI donated API credits for benchmarking their own evaluated models. The sample is modest at 110 tasks, and the authors chose not to report a controlled human baseline at all — citing cost (roughly 3 hours per task, an estimated $4,950 for all 110) — so there is no way to know what fraction a human completes for comparison. The live-web design means reproducibility is contingent on the state of the web at evaluation time, a limitation the authors acknowledge themselves. Results come from single runs with no repeated trials, so the variance behind these numbers is unknown.

Related Work

Compiled from the bibliographic details in this paper's own Related Work section and reference list. Session egress was blocked, so these three works could not be independently re-verified.

All three works form part of the same lineage of web and office agent benchmarks, but none, per the authors, jointly requires search plus artifact synthesis, formatting, and layout the way KNOWS does. Read alongside Long-Horizon Browser Agent Deployment Criteria, a pattern emerges across benchmarks: agents can search, but finishing the job is where they keep falling short.

Reviewer's Judgement

First, the most practically important finding here is not the 2.7% full-success rate but the gap between partial scores and real usability. An ACF/SF near 70% coexisting with an ASC of 0.14 on the showcased example means teams reading only aggregate partial-success metrics can easily conclude a task is "almost done" when it is not. Setting deployment thresholds on a single averaged metric is risky; a checkpoint-level completion metric like ASC needs to run alongside it.

Second, the finding that harness matters more than model (+18.7pp ASC for the same Opus 4.7) gives direction to teams building their own agent pipelines: investing in browsing infrastructure, action spaces, and tool interfaces may yield larger gains than upgrading to the latest model. Teams that defer harness work because a model leaderboard looks good should reconsider that ordering.

Third, the decision not to report any human baseline is a real gap. A bare estimate of "about 3 hours per task" gives readers little way to judge how bad 2.7% full success actually is relative to a human completing the same work, leaving that severity judgment entirely to the reader.

Putting It to Work

  • Track checkpoint-level completion alongside averages — pair ACF/SF with a stricter metric like ASC so partial-credit averages don't create a false sense of "almost done."
  • Re-rank harness investment against model upgrades — budget for browsing infrastructure, action spaces, and tool interfaces, which this study shows can outweigh swapping the underlying model.
  • Make visual input the default — accessibility-tree-only text access loses significant performance on tasks requiring visual or spatial judgment; default to a screenshot/VLM path.
  • Set deployment thresholds per artifact type — success rates and partial-score distributions differ across documents, spreadsheets, and slides, so don't apply one blanket bar across all of them.
  • Re-validate live-web-dependent tasks on a schedule — evaluation scripts tied to live websites can break as pages change, so give any internal eval set of this kind a periodic-check and deprecation policy.

Conclusion

KNOWS puts a number on something many teams already suspect: the step after search — synthesizing information into a finished document, spreadsheet, or slide deck — remains an open problem for today's frontier agents. Even the best pairing achieved only 2.7% full success, and cases with seemingly high partial scores still produced unusable artifacts. The finding that harness matters more than model has real implications for where agent-pipeline investment should go. That said, Google's and Anthropic/OpenAI's funding and credits, the absence of a human baseline, and the 110-task sample size all belong in how these numbers are read. The completion-rate problem in long-horizon tasks continues, with a different benchmark, in Long-Horizon Browser Agent Deployment Criteria.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…