arXiv:2601.08806cs.SEcs.AI2026-01被引 2

评测AI在真实软件工程任务中的生产力,发现强表现依赖严谨的推理与验证。

APEX-SWE

  • 构建两类真实工程任务:跨云系统集成与生产故障诊断。
  • Claude Opus 4.6以40.5%通过率领先,性能依赖事实与假设的区分能力。
  • 适合关注AI工程化落地、模型推理可信度的研究者与开发者。

我们提出软件工程领域的人工智能生产力指数(APEX-SWE),用于评估前沿AI模型执行具有经济价值的软件工程任务的能力。与聚焦单一、明确任务的现有评估不同,APEX-SWE引入两种反映真实场景的新任务类型:(1) 集成任务(共100项),需在异构云资源、业务应用和基础设施即代码服务间构建端到端系统;(2) 可观测性任务(共100项),需利用日志、仪表盘等遥测信号及非结构化上下文调试生产环境故障。我们对11个前沿模型进行了评估,Claude Opus 4.6以40.5%的Pass@1得分位居榜首,次为Claude Opus 4.5的38.7%。分析表明,优秀表现主要源于认知纪律——区分假设与已验证事实的能力,常伴随行动前的系统性验证。我们开源了APEX-SWE评估工具包及开发集(n=50)。

原文摘要 · Abstract (English)

We introduce the AI Productivity Index for Software Engineering (APEX-SWE), a benchmark for assessing whether frontier AI models can execute economically valuable software engineering work. Unlike existing evaluations that focus on narrow, well-defined tasks, APEX-SWE assesses two novel task types that reflect real-world software engineering: (1) Integration tasks (n=100), which require constructing end-to-end systems across heterogeneous cloud primitives, business applications, and infrastructure-as-code services, and (2) Observability tasks (n=100), which require debugging production failures using telemetry signals such as logs and dashboards, as well as unstructured context. We evaluated eleven frontier models for the APEX-SWE leaderboard. Claude Opus 4.6 leads the APEX-SWE leaderboard with 40.5% Pass@1, followed by Claude Opus 4.5 at 38.7%. Our analysis shows that strong performance is primarily driven by epistemic discipline, defined as the capacity to distinguish between assumptions and verified facts. It is often combined with systematic verification prior to acting. We open-source the APEX-SWE evaluation harness and a dev set (n=50).

AI工程评测基准软件开发推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。