Source
Daksh Raghuvanshi, Ved Vedere, Yifan Wang, "StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents", arXiv:2610.10942 [cs.AI], submitted 2026-10-07, DOI 10.48550/arXiv.2610.10942. License: CC BY 4.0. The authors released five training-split tasks, ten sample trajectories, and the scoring/verification code as an anonymized appendix, while withholding the simulation engine and the full evaluation suite to keep the benchmark uncontaminated.
This is an arXiv preprint that has not gone through peer review. The paper includes no section disclosing author affiliation or funding sources, and a WebSearch for the three authors (Daksh Raghuvanshi, Ved Vedere, Yifan Wang) turned up no independent confirmation of their affiliation — this gap is flagged plainly in the credibility assessment below. The 11 human experts who took part are described as "randomly selected from an existing pool of contract experts" and fairly compensated.
Study Overview
The research question boils down to one thing: can frontier LLMs, operating a simulated apparel store on real commerce software (the open-source platform Medusa v2) for 30 to 365 days, generate more profit and reliability than scripted heuristic policies and human experts?
The StoreBench environment hands the agent a store with 18 months of history already baked in (3,320 orders, roughly 2,500 customers, a 14.7% return rate), operable only through 29 MCP tools spanning catalog, pricing, inventory, fulfillment, returns, sourcing, and reporting. Time advances purely as a function of "operations spent," not inference latency, so a slower model is not penalized. Demand, supplier reliability, and competitor share are hidden variables the agent can only infer from its own sales data, and scripted events (demand shifts, supply failures, quality crises) are disclosed through three different channels — announced, rumored, or silent.
The evaluation has two tracks. A short-horizon suite (11 tasks × 3 seeds × 3 attempts, 30-45 days, under the goose harness) pits seven frontier models — DeepSeek-V4-Pro, Claude Opus 4.8, Claude Fable 5, Qwen3.8-Max, Gemini 3.8 Flash, GPT-5.6 Sol, Muse Spark 1.2 — against each other, while a separate full-year marathon task runs under the Claude Code harness. Every pass threshold is calibrated per task-seed pair against six scripted policies (do-nothing, absentee, blast-list, rule-based, smart-triage, exploiter) run through the identical tool surface.
Key Results
The composite score is a weighted sum of business performance (weight 0.7), service reliability (0.2), and continuity (0.1). Averaged over 11 tasks × 3 seeds × 3 attempts, no model matched the checklist-style scripted policy (smart-triage).
| Policy / Model | Composite | Pass | Business | Reliab. | Contin. |
|---|---|---|---|---|---|
| Smart-triage (heuristic) | 0.764 | 97% | – | – | – |
| DeepSeek-V4-Pro | 0.700 | 49% | 0.614 | 0.852 | 1.000 |
| Claude Opus 4.8 | 0.668 | 36% | 0.573 | 0.834 | 1.000 |
| Human experts (11) | 0.708 | 39% | 0.599 | 0.944 | 1.000 |
| Muse Spark 1.2 | 0.387 | 8% | 0.226 | 0.644 | 1.000 |
The best model, DeepSeek-V4-Pro (0.700, 49% pass rate), fell well short of smart-triage (0.764, 97%) and narrowly behind the mean human-expert score (0.708). Humans only led on the service side — 96% on-time fulfillment and 93% of returns handled, for a reliability score of 0.944 that beats every model — while their business component (0.599) actually trailed DeepSeek's (0.614). So "humans won" is only true on the reliability axis, not the business axis; the two bases should not be blended into a single verdict.
The ranking flips on the full-year marathon. DeepSeek (0.976), Qwen3.8-Max (0.974), Claude Opus 4.8 (0.971), and Claude Fable 5 (0.966) all cleared smart-triage (0.764) by a wide margin. But this comparison swaps both the horizon and the harness at once (goose vs. Claude Code, with context compaction and extended thinking) — the authors themselves state the two setups are "not a direct comparison." Rather than concluding "models overtake the heuristic long-term," the right reading is that a harness swap alone can move the result this much. Gemini 3.8 Flash was the one model whose score fell (0.552→0.496): in four of nine attempts it asserted a background simulation was running and simply waited (those four attempts averaged about 0.1).
A post-training experiment fine-tuned Qwen3.5-27B with GRPO on five tasks disjoint from the evaluation suite, then tested it on the 11 held-out evaluation tasks.
| Policy | Composite | Pass | Business | Reliab. | On-time | Returns |
|---|---|---|---|---|---|---|
| Base | 0.136 | 0/99 | 0.005 | 0.401 | 0.65 | 0.16 |
| Post-GRPO | 0.373 | 3/99 | 0.236 | 0.669 | 0.89 | 0.45 |
The composite rose from 0.136 to 0.373 (+0.237 per cell, bootstrap 95% CI [+0.19, +0.29]), improving 31 of 33 cells, and the share of profitable episodes jumped from 3% to 72%. Still, only 3 of 99 cells passed, and the rule-based checklist was never cleared on any task — the gains came from operating discipline (on-time 0.65→0.89, returns handled 0.16→0.45), not from pricing or sourcing judgment, which barely changed.
The most widespread failure mode was stockouts: 26% of the 693 recorded trials (178) under-restocked or restocked too late, and in 60 trials the agent never placed a single purchase order (41 of those were Gemini). The second is the native-tool-vs-shell split: every model could issue either native tool calls or batched shell scripts for identical access, but Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol, and Gemini 3.8 Flash drove the store through the shell in nearly every trial. The cost showed up within models: Muse scored 0.34 via shell versus 0.43 native, and Gemini scored 0.89 on its two native trials versus 0.55 on shell-driven ones (small sample sizes per model, so read this as directional only).
Credibility Assessment
Three things support trust here. First, every calibration measures six scripted policies run through the real tool surface for each task-seed pair, and a task is only accepted after clearing four gates (anchor ordering, a judgment gap, a noise margin, exploit suppression) — all 36 reported cells and 24 cells from two privately held-out seeds are stated to pass. Second, recomputing Table 1's composite formula (0.7×business + 0.2×reliability + 0.1×continuity) directly from the reported components matches the published scores to three decimal places — e.g. DeepSeek (0.7×0.614+0.2×0.852+0.1×1.0=0.700) and human experts (0.7×0.599+0.2×0.944+0.1×1.0=0.708). Third, 756 pre-hardening trials were audited for reward exploits, closed with regression tests, and the exploiter policy's score (0.039) is re-verified below every honest policy at every calibration pass.
The caveats are just as clear. This is an unreviewed preprint with no disclosed author affiliation, funding, or conflict-of-interest statement (see Source, above). The environment is a single fictional apparel store with scripted competitors and demand, so generalizing the results to other industries or real markets is a stretch — a limitation the authors acknowledge themselves. The human comparison uses one selected outcome per task-seed cell with one allowed retry, versus an average of three unselected model attempts, and humans were evaluated through the same restricted MCP interface as agents rather than their usual storefront tools — an asymmetry that could cut either way. That the authors flagged the harness mismatch in the long-horizon comparison is a transparency signal in their favor, but it's also exactly why "models beat humans over the long run" would be an overreach.
Related Work
- Backlund & Petersson (2025). Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents — a prior long-horizon coherence benchmark built on running a vending-machine business. StoreBench cites it directly as lacking a task-calibrated scale and pass threshold, and extends the paradigm onto a production-grade commerce backend (prior work, extension).
- Froger et al. (2026, ICLR). Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments — introduces scripted events independent of the agent, but caps its Time scenarios at five minutes; per a figure StoreBench's own text cites directly, removing generation latency raises GPT-5's score from 0% to 34.4%. StoreBench's metered-time design structurally removes this exact confound (prior work, methodological contrast; the number is verified directly from StoreBench's own primary text).
- Shi et al. (2026). MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations — another 365-day e-commerce operations benchmark; per independent WebSearch confirmation, its best model configuration reached only 27.3% of the mean human-participant net asset outcome (session egress was blocked, so the primary text could not be re-read; this figure is used only in the relation description below, not in Key Results). Set beside StoreBench's short-horizon result — where the best model reaches about 99% of the human mean — the human-model gap clearly depends heavily on benchmark design, which argues against generalizing any single benchmark's gap figure across the industry (related work, numeric contrast, benchmark origin).
Taken together, the question of whether agents can sustain long-horizon operations as well as humans has been asked repeatedly since Vending-Bench in 2025, and the answer shifts substantially with environment design choices — how time is metered, how calibration works, and how the human comparison is structured. StoreBench arguably has the strictest calibration and exploit-hardening procedure in this lineage, which makes the long-horizon reversal in Table 2 all the more in need of replication without a harness confound.
Reviewer's Assessment
First, the number that matters most here is not the 49% pass rate but the harness effect. The same models that fall short of a checklist policy on the short-horizon suite clear the heuristic's ceiling by a wide margin once the harness switches to Claude Code. That the authors flagged this as "not a direct comparison" upfront is unusually honest, but it also demonstrates how risky it is for benchmark leaderboards to report "model X scored Y" without naming the harness.
Second, the GRPO result (0.136→0.373) should not be read as evidence that training solves autonomous operation. Only 3 of 99 cells passed, the rule-based checklist was never cleared, and the gains were confined to operating discipline (on-time shipping, returns) rather than pricing or sourcing judgment. That five training tasks transfer any signal to held-out tasks is an interesting sample-efficiency finding, but it's premature as grounds for a product decision.
Third, that three of the four weakest-performing closed models default to shell access over native tool calls shows that "the agent had MCP tools available" and "the agent actually used them" are different claims. The authors' own observation that shell-driven operation correlates with batch decisions made on stale state suggests that leaving shell access open in an agent operations system carries a real cost of its own.
Practical Takeaways
- Benchmark against a scripted heuristic first — before deploying an operations agent, measure it against a checklist-level scripted policy. By this paper's numbers, even frontier models don't clear that bar.
- Evaluate with the harness you'll actually deploy — the same model's score can swing drastically with tool-exposure and context-management choices, so re-test under the exact production harness.
- Track stockouts and under-refunds as separate KPIs — stockouts were the dominant failure mode at 26% of trials, so monitor missed purchase orders and under-refunding on the operations dashboard directly.
- Audit native-tool-call usage — log whether the agent is actually calling the tools it's given, or falling back to shell access and losing track of state.
- Treat small-scale RL fine-tuning as a supporting measure, not a fix — a handful of training tasks can improve operating discipline without transferring commercial judgment, so don't hand pricing or sourcing decisions to a trained policy without separate validation.
Conclusion
StoreBench's contribution is the first calibrated, exploit-hardened answer to "can an agent run a store?" The answer isn't simple: no frontier model cleared the scripted checklist on the short-horizon suite, yet changing the harness and extending the horizon flips the ranking entirely. That inconsistency is itself the practical lesson — agent performance is a property of the model-harness-horizon combination, not the model alone, so pre-deployment validation needs to test that whole combination. A companion piece on detecting and escalating anomaly signals when running agent evaluation environments is available at An Evaluation-Environment Anomaly Escalation Gate.
References
- StoreBench — arXiv abstract (primary source)
- Same paper, HTML full text — used for table/number verification
- Vending-Bench (Backlund & Petersson, 2025) — related work source
- Gaia2 (Froger et al., 2026) — related work source
- MerchantBench (Shi et al., 2026) — related work source
- An Evaluation-Environment Anomaly Escalation Gate — sunny34.com blog
Ask AI about this review
The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…