arXiv:2512.23844cs.SEcs.AI2025-12被引 2

AI写代码要能协作,不能只看对错。

From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering

  • 从91组用户规则提炼出四类核心协作行为。
  • 发现评价标准随时间与任务类型动态变化。
  • 适合关注AI人机协作的工程师和研究者。

随着大语言模型从代码生成工具演变为软件工程师的协作伙伴,当前的评估方法已落后。现有基准主要关注代码正确性,无法捕捉人类与AI协作所需的复杂互动行为。本文提出两项核心贡献:首先,基于对91组用户定义的代理规则分析,构建了企业级软件工程中理想代理行为的基础分类体系,涵盖四大期望:遵循规范与流程、保障代码质量与可靠性、有效解决问题、与用户协同工作;其次,认识到这些期望并非固定不变,提出了情境自适应行为(CAB)框架。该框架通过15位专家访谈确定时间维度(从即时需求到未来理想),并结合原型代理的提示分析识别任务类型维度(如企业生产与快速原型开发)。两项贡献共同为下一代AI代理的设计与评估提供了以人为本的基础,推动领域从代码正确性转向真实协作智能的动态评估。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) evolve from code generators into collaborative partners for software engineers, our methods for evaluation are lagging. Current benchmarks, focused on code correctness, fail to capture the nuanced, interactive behaviors essential for successful human-AI partnership. To bridge this evaluation gap, this paper makes two core contributions. First, we present a foundational taxonomy of desirable agent behaviors for enterprise software engineering, derived from an analysis of 91 sets of user-defined agent rules. This taxonomy defines four key expectations of agent behavior: Adhere to Standards and Processes, Ensure Code Quality and Reliability, Solving Problems Effectively, and Collaborating with the User. Second, recognizing that these expectations are not static, we introduce the Context-Adaptive Behavior (CAB) Framework. This emerging framework reveals how behavioral expectations shift along two empirically-derived axes: the Time Horizon (from immediate needs to future ideals), established through interviews with 15 expert engineers, and the Type of Work (from enterprise production to rapid prototyping, for example), identified through a prompt analysis of a prototyping agent. Together, these contributions offer a human-centered foundation for designing and evaluating the next generation of AI agents, moving the field's focus from the correctness of generated code toward the dynamics of true collaborative intelligence.

人机协作AI代理评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。