SWE-Bench Pro构建了更复杂的软件工程任务基准,测试AI能否完成企业级长期开发任务。
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- 基于41个活跃仓库构建1865个真实复杂任务,涵盖多文件修改与长时间开发
- 包含公开、保留与商业三类数据集,其中18个为初创企业专有代码库
- 聚焦长期任务与错误模式分析,助力评估AI代理在专业场景下的真实能力
我们提出SWE-Bench Pro,一个比SWE-BENCH更具挑战性的基准,旨在捕捉超出原基准范围的真实复杂企业级问题。该基准包含来自41个活跃维护仓库的1,865个问题,涵盖业务应用、B2B服务和开发者工具。基准分为三部分:11个仓库的公开集、12个仓库的保留集,以及18个拥有正式合作的专有仓库商业集。保留集与商业集不公开,但我们会发布商业集上的评测结果。所有任务均需数小时至数天完成,常涉及跨多个文件的代码修改。每个任务经人工验证并附充分上下文以确保可解性。我们对现有模型轨迹中的失败模式进行聚类分析,以清晰刻画当前模型的错误特征。SWE-Bench Pro提供抗污染的测试环境,更真实地反映现实软件开发的复杂性与多样性,推动实现专业级自主软件工程智能体的发展。
原文摘要 · Abstract (English)
We introduce SWE-Bench Pro, a substantially more challenging benchmark that builds upon the best practices of SWE-BENCH [25], but is explicitly designed to capture realistic, complex, enterprise-level problems beyond the scope of SWE-BENCH. SWE-BENCH PRO contains 1,865 problems sourced from a diverse set of 41 actively maintained repositories spanning business applications, B2B services, and developer tools. The benchmark is partitioned into a public set with open access to problems sourced from 11 repositories, a held-out set of 12 repositories and a commercial set of 18 proprietary repositories where we have formal partnership agreements with early-stage startups. Problems in the held-out and the commercial set are not publicly accessible, but we release results on the commercial set. Our benchmark features long-horizon tasks that may require hours to days for a professional software engineer to complete, often involving patches across multiple files and substantial code modifications. All tasks are human-verified and augmented with sufficient context to ensure resolvability. To better understand these limitations, we cluster the failure modes observed in the collected agent trajectories for a clearer characterization of the error patterns exhibited by current models. Overall, SWE-BENCH PRO provides a contamination-resistant testbed that more faithfully captures the complexity and diversity of real-world software development, advancing the pursuit of truly autonomous software engineering agents at a professional level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。