构建对话式编程评测基准,测试代码助手真实交互能力
Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents

- 设计角色驱动的用户模拟器,模拟真实开发对话场景
- 新模型在对话质量上提升3-14%,超越现有基线
- 发现代码能力与对话能力不相关,需独立评估
AI编程助手已广泛应用于实际开发,但现有评测将它们视为完全自主系统。本文提出Dialogue SWE-Bench,一个面向对话式编程任务的自动评测数据集,用于评估编码代理通过与用户对话解决真实软件工程问题的能力。我们设计了一种基于角色的用户模拟器,并引入自动对话质量评估机制。此外,提出一种基于模式引导的新型代理架构,显著提升现成编码代理的对话能力,在多个指标上优于强基线3%-14%。实验表明,更强的代码模型并不一定具备更好的对话能力,说明对话能力是编码代理性能中一个独立且尚未被充分研究的维度。
原文摘要 · Abstract (English)
AI coding agents have rapidly transformed software engineering, powering widely used interactive coding assistants. Despite their interactive real-world use, existing benchmarks evaluate them as fully-autonomous systems. In this work, we introduce Dialogue SWE-Bench, an automatic benchmark dataset for evaluating the ability of coding agents to resolve real-world software engineering problems through dialogue with a user. We design a novel, persona-grounded user simulator to support our task evaluation, and augment our task evaluation with automatic evaluations of dialogue quality. We also propose a new schema-guided agent, aimed at improving the dialogue capabilities of off-the-shelf coding agents, which improves over strong baselines by 3-14%. Our results indicate that better coding models do not always correspond to better dialogue models, suggesting that dialogue capability is a distinct and currently understudied dimension of coding agent performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。