arXiv:2603.13428cs.SEcs.AI2026-03被引 6

新基准评估AI代码代理在持续演化中的长期维护能力。

SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution

  • 构建可验证的里程碑任务图,模拟真实软件演进过程。
  • 12个前沿模型在连续任务中表现下降至38.03%。
  • 适合研究长期代码生成与系统稳定性的团队。

现实世界软件需持续演进以应对不断变化的开放需求。日益部署为长时运行系统的AI代理正被赋予驱动这一演进的任务。然而,现有基准仅评估代理在孤立、单次编码任务上的表现,忽视了真实软件演进中的时间依赖性和技术债务。为此,我们提出DeepCommit,一种从噪声提交日志重建可验证里程碑有向无环图(Milestone DAG)的智能体流水线,其中里程碑定义为功能一致的开发目标。这些可执行序列构成SWE-Milestone基准,用于评估代理在里程碑级任务流上的表现,要求其维持系统完整性并控制错误累积——这正是当前基准普遍缺失的长期软件演进维度。我们在4个代理框架上对12个前沿模型的评估显示:整体性能得分从孤立任务的>80%显著下降至连续设置下的38.03%,暴露出代理在长期维护和错误传播方面存在根本性脆弱性。

原文摘要 · Abstract (English)

Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-running systems, are now entrusted to drive this evolution. Yet, existing benchmarks evaluate agents on isolated, one-off coding tasks, neglecting the temporal dependencies and technical debt inherent in real-world software evolution. To bridge this gap, we introduce DeepCommit, an agentic pipeline that reconstructs verifiable Milestone DAGs from noisy commit logs, where milestones are defined as functionally cohesive development goals. These executable sequences enable SWE-Milestone, a benchmark that evaluates agents on streams of milestone-level tasks, requiring them to sustain system integrity and limit error accumulation, dimensions of long-term software evolution largely missing from current benchmarks. Our evaluation of 12 frontier models across 4 agent frameworks reveals a critical vulnerability: overall performance scores drop significantly from >80% on isolated tasks to 38.03% in continuous settings, exposing agents' profound struggle with long-term maintenance and error propagation.

AI代理代码生成持续演进基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。