评测代码智能体主动发现并修复多个漏洞,无需问题报告。
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

- 提出新任务范式与双轨评估框架,支持主动查错
- 覆盖1663个任务,涵盖8种语言和6类漏洞
- 揭示主流模型在主动修复上表现有限,适合研究者参考
由大语言模型驱动的代码智能体在软件工程中日益普及,能够修复大规模代码库中的特定漏洞。然而,现有软件工程基准通常假设高质量的问题报告始终可用,这在实践中因报告获取与整理复杂而难以满足。为此,我们提出了 Active-SWE,一个用于评估代码智能体在无报告指引下主动发现并修复多个漏洞的基准,涵盖1,663个任务,覆盖六类漏洞和八种编程语言。该基准不仅将评估重点从被动修复转向主动发现,还扩展了评估范围,从修复已知漏洞到多漏洞修复及潜在漏洞发现。为构建 Active-SWE,我们提出一种新颖的难度感知任务生成流程与双轨评估框架,全面评估主动修复能力。大量实验表明,大多数顶尖代码智能体在主动修复任务中表现不佳,其定位与修复已知漏洞、处理多漏洞场景及发现有效潜在漏洞的能力均有限。
原文摘要 · Abstract (English)
Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation. To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages. Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios. To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability. Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potential bugs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。