量化关键信息对代码生成代理性能的贡献,指导自主编程系统研发。
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents

- 构建统一方法提取代码任务中的关键信息信号
- 发现理想状态下各信号可提升代理成功率超40%
- 适合研究自主编程系统与智能助手的开发者
语言模型代理在自动化软件工程领域取得显著进展。现有研究已探索多种代理工作流、训练策略及失败模式,关注包括复现测试、回归测试、编辑位置、执行上下文和API使用在内的多种上下文信息信号。然而,各信号对整体成功贡献的具体影响仍不明确,尤其缺乏在中间信息完美获取的理想情境下的评估。为此,本文提出Oracle-SWE,一种统一方法,用于从SWE基准中分离并提取这些理想信息信号,量化其对代理性能的影响。为进一步验证规律,我们评估强语言模型提取的信号提供给基础代理时带来的性能提升,模拟真实任务解决场景。该研究旨在为自主编程系统的研发方向提供指导。
原文摘要 · Abstract (English)
Recent advances in language model (LM) agents have significantly improved automated software engineering (SWE). Prior work has proposed various agentic workflows and training strategies as well as analyzed failure modes of agentic systems on SWE tasks, focusing on several contextual information signals: Reproduction Test, Regression Test, Edit Location, Execution Context, and API Usage. However, the individual contribution of each signal to overall success remains underexplored, particularly their ideal contribution when intermediate information is perfectly obtained. To address this gap, we introduce Oracle-SWE, a unified method to isolate and extract oracle information signals from SWE benchmarks and quantify the impact of each signal on agent performance. To further validate the pattern, we evaluate the performance gain of signals extracted by strong LMs when provided to a base agent, approximating real-world task-resolution settings. These evaluations aim to guide research prioritization for autonomous coding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。