让编程助手学会揣摩用户心思,提升代码协作效率
TOM-SWE: User Mental Modeling For Software Engineering Agents
- 用双代理架构,让助手理解用户意图和偏好
- 在状态化任务中成功率从18.1%提升至59.7%
- 适合需要长期协作的开发者日常使用
近年来编码代理已能完成复杂代码库的规划、编辑、运行与测试。然而,它们在推断和追踪用户意图方面仍存在困难,尤其当指令不明确或依赖上下文时。为此,我们提出ToM-SWE,一种由主软件工程(SWE)代理与轻量级心智模型(ToM)伙伴代理组成的双代理架构。ToM代理从指令和交互历史中推断用户目标、约束与偏好,维护用户的持续记忆,并向SWE代理提供相关建议。在两个软件工程基准测试(模糊SWE-bench与状态化SWE-bench)中,ToM-SWE显著提升了任务成功率与用户满意度。特别是在新引入的状态化SWE基准上,结合用户模拟器与历史交互记录,ToM-SWE达到59.7%的任务成功率,远超当前顶尖的OpenHands代理(18.1%)。此外,在为期三周的专业开发者实测中,参与者在86%的时间认为该系统有用,证明了持续用户建模对实际编程代理的价值。
原文摘要 · Abstract (English)
Recent advances in coding agents have made them capable of planning, editing, running, and testing complex code bases. Despite their growing ability in coding tasks, these systems still struggle to infer and track user intent, especially when instructions are underspecified or context-dependent. To bridge this gap, we introduce ToM-SWE, a dual-agent architecture that pairs a primary software-engineering (SWE) agent with a lightweight theory-of-mind (ToM) partner agent dedicated to modeling the user's mental state. The ToM agent infers user goals, constraints, and preferences from instructions and interaction history, maintains a \textbf{persistent memory} of the user, and provides user-related suggestions to the SWE agent. In two software engineering benchmarks (ambiguous SWE-bench and stateful SWE-bench), ToM-SWE improves task success rates and user satisfaction. Notably, on the stateful SWE benchmark, a newly introduced evaluation that provides agents with a user simulator along with previous interaction histories, ToM-SWE achieves a substantially higher task success rate of 59.7\% compared to 18.1\% for OpenHands, a state-of-the-art SWE agent. Furthermore, in a three-week study with professional developers using ToM-SWE in their daily work, participants found it useful 86\% of the time, underscoring the value of stateful user modeling for practical coding agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。