Source
Merve Astekin, Yan Naing Tun, Arda Goknil, Erik Johannes Husom, Lwin Khin Shar, Hasan Sözer, Ratnadira Widyasari, Hui Song, "Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows", arXiv:2610.03010 [cs.SE], submitted 2026-10-02 (v1), DOI 10.48550/arXiv.2610.03010. Affiliations: SINTEF (Norway), Singapore Management University, Özyeğin University (Türkiye). The full text was cross-checked against a primary-source snapshot taken 2026-10-05T22:14:44Z (this session's network egress was fully blocked — confirmed via three retries at 10-minute intervals before falling back).
This is an unreviewed preprint. Funding is explicitly stated as EU Horizon Europe grants ENFIELD (101120657) and INTEND (101135576); no LLM-vendor funding or author/evaluated-model conflict of interest was found — all six evaluated models are open-weight public releases.
Study Overview
Two questions drive the study: how do accuracy, energy, and latency change as agent count increases, and does that trade-off point the same way across tasks? The authors test four configurations — non-agentic single-query (NA), single-agent (SA), dual-agent (DA), and four-agent multi-agent (MA) — across five software-engineering tasks (code generation via HumanEval-style tasks, technical-debt detection on MLCQ/Java, vulnerability detection on PrimeVul/C-C++, log parsing, and log analysis on HDFS), six open-weight models (up to 20B parameters), two prompting strategies (zero-shot/few-shot), and three hardware platforms (two servers, one workstation), with repeated runs. Metrics are task-specific accuracy (Pass@1, Macro F1, F1, or exact-match, depending on the task), end-to-end latency, and total pipeline energy consumption (kWh).
Key Results
Averaged across all 15 task–hardware pairs, models, and prompts (Tables 16 and 20):
| Configuration | Mean energy (kWh) | vs. NA | Mean latency (s) | vs. NA |
|---|---|---|---|---|
| NA (non-agentic) | 0.128 | 1.00× | 1,789.5 | 1.00× |
| SA (single-agent) | 0.154 | 1.20× | 2,205.9 | 1.23× |
| DA (dual-agent) | 0.392 | 3.06× | 5,940.0 | 3.32× |
| MA (four-agent) | 0.813 | 6.36× | 10,865.5 | 6.07× |
The worst case is log parsing on Workstation-SMU, where latency goes from 313.1s (NA) to 50,185.8s (MA, nearly 14 hours) — a 160× slowdown. Accuracy, however, did not move in the same direction as energy and latency.
| Task (metric) | NA | SA | DA | MA |
|---|---|---|---|---|
| Vulnerability detection (F1) | 49.7% | 49.2% | 35.6% | 55.4% |
| Log parsing (exact-match) | 34.5% | 31.2% | 27.6% | 26.1% |
The other three tasks stayed in a narrow band across configurations: code generation Pass@1 94.2–96.8%, technical-debt detection Macro F1 40.0–42.8%, and log analysis F1 5.5–7.5% (uniformly low because of class imbalance — anomalous sessions are only about 3% of the HDFS dataset). Vulnerability detection is the only task where MA beats every other configuration. Across the three-objective (accuracy, energy, latency) Pareto front of 66 optimal configurations over 15 task–hardware pairs: NA accounts for 36 (55%), SA for 23 (35%), DA for 6 (9%), and MA for just 1 (2%) — NA+SA together make up 59 (89%).
Credibility Assessment
Reasons to trust it: the work is publicly funded (EU Horizon Europe) with no vendor conflict of interest, and all evaluated models are open-weight, so the setup is reproducible. Temperature was fixed at 0, most configurations were run three times, statistical significance was tested, and RQ1's accuracy analysis balances hardware coverage across models to avoid skew.
Caveats: this is an unreviewed preprint. The most important limitation is that every evaluated model is a small, locally hosted open-weight model (≤20B parameters) — the energy and latency multipliers may not transfer directly to frontier hosted-API models (GPT-5-class, Claude, Gemini). GPT-OSS was run only once, on a single hardware platform, with no repetitions. The datasets are limited to Python, Java, and C/C++, and are curated benchmarks rather than noisy real-world repositories. As conflicting evidence, Shu et al. (2024) report up to a 70% success-rate gain from multi-agent over single-agent in a different comparison — the discrepancy likely reflects differing task types (enterprise coordination/routing vs. the five SE tasks here) and whether cost was measured at all.
Related Work
- Shu et al. (2024), "Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications" — Contradicting/complementary (AWS Bedrock). Reports up to a 70% success-rate gain from multi-agent over single-agent on enterprise coordination/routing tasks, without measuring energy or latency cost. The different task domain is compatible with this paper's finding that multi-agent gains are task-specific.
- Husom et al. (2024), "The Price of Prompting: Profiling Energy Use in Large Language Models Inference" — Prior work (overlapping authors; submitted to NeurIPS 2024). Introduces the MELODI framework and finds 70B-class models use up to two orders of magnitude more energy per token than smaller models for single-model inference. This paper extends the same energy-profiling approach from single-model inference to full multi-agent pipelines.
- Alizadeh et al. (2025), "Language Models in Software Development Tasks: An Experimental Analysis of Energy and Accuracy" — MSR 2025 — Prior/extension. Shows that bigger models and higher energy budgets don't reliably improve accuracy, and no single model fits every SE task. This paper reconfirms the same "spending more doesn't mean better" pattern along the axis of agentic architecture complexity rather than model size.
All three were independently identified and confirmed via web search (full-text re-access was blocked by this session's egress policy); their titles, authors, and venues match this paper's own reference list.
Reviewer's Take
First, I don't think "is the task reasoning-heavy?" is the right test for whether to add agents. The one task where MA wins — vulnerability detection — is one where cross-checking from different angles reduces failures; code generation and log analysis involve plenty of reasoning but no need to split perspectives, and multi-agent lost there. "Does cross-verification reduce failures?" strikes me as the sharper, more actionable test.
Second, I think the most underrated result is that log-parsing accuracy fell as agents were added, from 34.5% to 26.1%. For format-sensitive tasks, the agent-to-agent handoff itself may be corrupting output structure. Read alongside the authors' own G3 finding — few-shot prompting alone raised log-parsing accuracy from 3.4% to 54.3% — this reads as evidence to fix the prompt before scaling the architecture.
Practical Takeaways
- Start non-agentic, escalate only with evidence — NA was best on at least one of accuracy/energy/latency for all five tasks. Treat adding agents as a hypothesis to test, not a default.
- Fix the model before the architecture — the gap between the best and worst model choice (up to two orders of magnitude in energy) exceeds the gap between agent configurations. Check model-task fit before scaling up the pipeline.
- Use few-shot for format-sensitive tasks — for tasks like log parsing, few-shot prompting is cheaper and more effective than adding agents (+50.9pp accuracy, 24–41% less energy in this study).
- Reserve hardware upgrades for heavy workloads — for DA/MA, server-class hardware is up to 3.4× faster than a workstation. For NA/SA, hardware differences are small enough to pick on availability alone.
- Report energy alongside accuracy — publishing accuracy only hides this kind of trade-off. Measure energy and latency on a fixed reference platform.
Conclusion
This study shows, with a controlled measurement across five tasks and three hardware platforms, that multi-agent architecture should be treated as a hypothesis, not a default. On average, the four-agent configuration used 6.36× the energy and 6.07× the latency of the non-agentic baseline, yet improved accuracy on only one of five tasks (vulnerability detection), and 89% of Pareto-optimal configurations were non-agentic or single-agent.
The caveat is that every evaluated model is a small, locally hosted open-weight model — teams running frontier hosted APIs should re-verify these multipliers on their own stack rather than importing them directly. Designing a halting rule before adding agents is the operational follow-through covered in One Stopping Rule Cut Cost by 63%: An Ops Guide to Multi-Agent Memory Gating.
References
- Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows — arXiv abstract (primary source)
- Full HTML text of the same paper — used for table/energy/latency cross-checks
- One Stopping Rule Cut Cost by 63%: An Ops Guide to Multi-Agent Memory Gating — sunny34.com blog
Ask AI about this review
The assistant has read this review and the numbers it verified. Ask anything — it answers from the text and says so when something isn't in it.
Loading the chat…