Source Document

Panagiotis Kasnesis, Christos Chatzigeorgiou, Lazaros Toumanidis, Amalia Contiero Syropoulou, "Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System", arXiv:2610.09021 [cs.AI], submitted 2026-10-06, DOI 10.48550/arXiv.2610.09021. Affiliation: University of West Attica (Athens, Greece), Waldiez PC, ThinGenious PC. Accepted at the NeurIPS 2026 Workshop on SLMs for Agentic Systems (Paris). Benchmark, harness and all 2,520 records released (github.com/waldiez/slm-callsite-eval).

This is a workshop paper, not a main-track peer-reviewed one — workshop bars are typically lighter. Session egress was blocked, so the full text was checked against the arXiv HTML snapshot fetch_source_snapshot.py collected at 2026-10-08T22:09:49Z. The conflict of interest is explicit: three of the four authors are at Waldiez PC, the company behind Wactorz, the open-source framework this paper evaluates; the fourth is at ThinGenious PC. The team that built and deployed the system is also the team grading it. Funding is disclosed as the NVIDIA Academic Grant Program (two donated DGX Spark systems), which can overlap with a conclusion that favors local inference economics.

Study Overview

The question is simple: an agentic system issues LLM calls of wildly different difficulty, yet practice usually serves them all with one model sized for the hardest call. Is that the right default? Rather than simulating, the authors measure a deployed open-source home-automation multi-agent framework, Wactorz (long-lived agents supervised over MQTT), and treat per-call-site model assignment as a configuration value rather than a runtime router.

The five call sites are: (1) intent routing (a four-way label under a ten-token budget), (2) Home Assistant action classification, (3) actuator — mapping language to JSON service calls against the live entity registry, (4) pipeline planning, and (5) code generation for the agent code that plan names. The first three use the framework's unmodified production prompts; the two generative sites use harness variants that preserve the output contract. Nine models (Qwen3.5 0.8B/2B/4B, Gemma4 E2B/E4B/12B/26B, Claude Haiku 4.5, Claude Sonnet 5) were compared across two real Home Assistant installations (a household and an office), 280 cases, 2,520 scored calls.

Key Results

The table below shows per-site accuracy and hosted cost across the full benchmark (n=280; local models cost nothing per call). Orderings differ by site and aren't monotonic in size — Qwen3.5 4B scores below the 2B model on actuation (50.9% vs. 62.1%), and Gemma4 26B falls below the 12B model on code generation (80.0% vs. 84.0%).

ModelIntentHA actionActuatorPlannerCodegenAllCost($)
Qwen3.5 0.8B50.071.433.68.024.033.20
Qwen3.5 4B91.792.950.952.038.058.20
Gemma4 26B91.796.494.888.080.090.70
Claude Haiku 4.580.692.993.182.094.089.63.06
Claude Sonnet 594.492.999.184.0100.095.412.13

Paired McNemar tests on the same cases show the best local model (Gemma4 26B) statistically indistinguishable from both hosted models at four of five sites. Code generation is the exception: p=0.039 against the small hosted model (Haiku 4.5), p=0.002 against the frontier one (Sonnet 5) — since a cheap hosted model reproduces the gap too, the authors read it as a property of the call site, not of frontier scale. Actuator misses significance against Sonnet 5 (p=0.062), but all five discordant cases point the same direction, which the authors read as insufficient statistical power rather than a true null.

Turned into routing policy: sending every call site to its best local model reaches 91.8% accuracy at zero per-call cost (vs. 95.4% for Sonnet 5 alone, -3.6pp). Hosting only the two generative sites (planner, codegen) reaches 93.9% (-1.4pp) while cutting spend from $12.13 to $2.15 — 82% less. 81% of hosted spend goes to the actuator call (its prompt carries the full entity registry, averaging 28,122 input tokens), yet that is exactly the site a mid-sized local model (Gemma4 12B, 95.7%) already handles best — the call that dominates spend isn't the call that needs the most capability.

A live-deployment comparison (46 cases, 43 scored under all three configurations; 3 excluded for framework-level issues unrelated to the model) reproduces the same trade-off.

ConfigurationDirect (26)Pipeline (20)Spend($)
All hosted (Sonnet 5)25/2615/202.47
Generative sites hosted24/2615/200.70
All local24/2612/200.00

Hosting only the generative sites matched all-hosted at 39/43, for 28% of the spend ($0.70 vs. $2.47). Aggregate accuracy hides the safety side, though: at the actuator site, Gemma4 E2B issued service calls on 87.2% of requests for devices the installation doesn't own, while Qwen3.5 0.8B returned an empty array on every request — its 116 "refusals" that score as correct reflect a model that recognized nothing at all, not calibrated judgment.

Credibility Assessment

Three things support trust: the core comparison runs all nine models on the same 280 cases under identical conditions (temperature 0, 4-bit quantization, one host); the offline benchmark's prediction (a -4.3pp actuator gap moving from Sonnet 5 to Gemma4 26B) nearly matched the live-deployment gap (1 case in 26, -4pp); and the authors disclosed three label revisions and their effect (+0.08 for the hosted model, -0.03 for local ones — against their own conclusion) rather than hiding them.

Caveats are substantial. The largest is conflict of interest — researchers at the company that built the evaluated framework graded their own system, with NVIDIA hardware support. The study runs one pass per case with no variance reporting, though the authors note one live case flipped outcome across configurations. The live-deployment judge was a single, unblinded rater. Only two installations were tested, so generalization to other layouts and device mixes is unverified. As contrasting evidence, Belcak et al. (below) argue most small-model calls suffice — broadly supported here, but directly contradicted at the code-generation site.

Related Work

Together, these three suggest small-model substitution has moved past the position-paper stage into deployed evidence, but calibrating valid from invalid requests — especially at generative and safety-relevant call sites — keeps resurfacing regardless of model or benchmark.

Reviewer's Judgement

First, I'd argue the most important finding isn't 91.8% or the 82% savings, but that aggregate accuracy hides a safety failure. Gemma4 E2B and Qwen3.5 0.8B can post similar overall actuator scores while one acts on 87.2% of requests for devices it doesn't own and the other refuses everything without recognizing anything. Any call site with irreversible physical effects — a lock, a thermostat — needs acting-rate and refusal-rate reported separately by default, not folded into one accuracy number.

Second, "host only the generative sites" looks like the most attractive trade-off at -1.4 points for -82% cost, but it's validated on one framework and two installations, so it shouldn't be copied as a number. What transfers is the procedure — split your own agent's call sites, measure each independently, then design routing from that.

Putting It to Work

  • Assign models per call site, not per system — a static configuration value set when you already know which call is being made, no runtime router needed.
  • Separate acting-rate from refusal-rate for actuation-type sites — a single accuracy number can't distinguish over-acting from blanket refusal.
  • Track spend by token, by call site — here, one prompt carrying the full device registry drove 81% of hosted spend. The costliest call site isn't necessarily the hardest one.
  • Deprioritize moving framework-specific code generation to local models — unfamiliar async callback and lifecycle API shapes are where the local-vs-hosted gap persisted longest.
  • Reweight offline benchmark scores by real call frequency — benchmarks weight categories equally, but a deployment's call mix doesn't, so offline accuracy shouldn't be read as a direct forecast of production performance.

Conclusion

This paper's contribution isn't a new model but a reframing: model selection should be a per-call-site decision, not a whole-system procurement choice. Local models were statistically indistinguishable from hosted ones at four of five sites, and hosting only the two generative sites held up in live deployment too, cutting cost 82% while keeping success nearly flat. What remains unresolved is that the team evaluating the system built it, variance isn't reported, and only two installations were tested. Take the procedure — split call sites, measure them independently, score safety-relevant sites separately — rather than the specific numbers. For the operational angle on cost optimization, see Cutting Cost 10x with Small-Model Task Routing.

References

Ask AI about this review

The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.

Loading the chat…