提出编码代理主动性的评估框架,区分主动与自主。
Agentic Coding Needs Proactivity, Not Just Autonomy

- 构建三层次主动性分类:反应式、计划式、情境感知式。
- 提出洞察决策质量等三项可量化的评估指标。
- 适合研究智能编程助手与人机协作的学者和开发者。
编码代理正从代码补全演变为能自主编辑仓库、创建拉取请求、响应问题并触发定时或事件驱动任务的系统。下一代将更强调主动性和长时程能力:在开发者提问前发现相关变更,跨工具关联信号,判断何时打断,以及保持会话间偏好。然而,当前领域仍缺乏对主动性的清晰定义,未明确其与自主性的区别,也无明确的接受标准及衡量无故行为是否有效的指标。本文主张以洞察策略的质量和改进程度来评估主动编码代理,该策略决定下一步关注什么、证据支持为何、是否展示结果及如何根据反馈调整。基于混合协作原则,提出三级主动性分类(反应式、计划式、情境感知式),对比现有编码代理的五项实践标准,并设计一种用户模拟评估协议,包含三个目标:洞察决策质量(IDQ)、上下文锚定得分(CGS)和学习提升度。
原文摘要 · Abstract (English)
Coding agents are rapidly changing the landscape of software development, moving from inline completion to autonomous systems that edit repositories, open pull requests, respond to issues, and run scheduled or webhook triggered routines across the development life cycle. The next generation is increasingly described as proactive and long-horizon: agents should notice relevant changes before the developer asks, connect signals across tools, decide when to interrupt, and carry preferences across sessions. Yet the field still lacks a clear account of what proactivity means for software development, how it differs from autonomy, what acceptance criteria proactive long-horizon tasks should satisfy, and which metrics determine whether unsolicited agent behavior is useful rather than merely active. Proactive coding agents should be evaluated by the quality and improvement of their insight policy: the policy that decides what matters next, what evidence supports it, whether to show it, and how to adapt after feedback. This view is grounded in the principles of mixed initiative interaction. We propose a three level taxonomy of proactivity (Reactive, Scheduled, and Situation Aware), compare contemporary coding agents against five practical criteria, and sketch an active user simulation protocol with three evaluation targets: Insight Decision Quality (IDQ), Context Grounding Score (CGS), and Learning Lift
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。