arXiv:2606.17799cs.SEcs.AI2026-06

现有编程基准无法真实反映智能体开发系统的表现

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

论文配图:Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
图 1 · 摘自论文原文
  • 将模型、工具链和环境合并为单一评分,掩盖各组件贡献
  • 单参考解导致合理替代方案被错误惩罚
  • 缺乏组件级反馈,难以优化系统迭代

编程智能体已成为软件工程主流模式,但当前评测基准诞生于代理时代之前:它们将模型、工具链和环境融合为单一端到端分数,通常仅与单一参考解对比,且缺少对迭代过程的分量信号。我们指出,现有基准与智能体式软件工程严重脱节。实际中的编程智能体并非单一模型,而是由模型、工具链、上下文、运行环境和反馈信号构成的复合系统,其中任一组件的改进都可能带来与相邻模型代际相当的分数提升。我们提出三个症状:(i) 基准分数混淆了模型与工具链的贡献;(ii) 以单一参考解为标准,惩罚了同样有效的替代方案;(iii) 缺乏对工具链各组件的独立信号,使整体系统得分难以有效迭代。

原文摘要 · Abstract (English)

Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era: they collapse model, harness, and environment into a single end-to-end score, typically computed against one reference solution, with no component-level signal for iteration. We argue that current coding benchmarks are misaligned with agentic software engineering. A coding agent in practice is not a model: it is a system harness -- a composite of models, harnesses, contexts, environments, and feedback signals, any one of which can move the benchmark score by margins comparable to those between adjacent model generations. We discuss three symptoms: (i) benchmark scores conflate the model with the rest of the harness; (ii) grading against a single reference solution penalises equally valid alternatives; and (iii) the absence of signal at the level of individual harness components makes the end-to-end system score difficult to iterate on.

编程智能体评测基准系统优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。