为编程大模型代理提供真实开发场景下的可靠评估方法
Reliable and Developer-Aligned Evaluation of Agents for Software Engineering
- 基于真实开发环境设计评估框架,避免虚构语法测试
- 关注代理在复杂项目中的行为轨迹与失败模式
- 适合开发者和研究者用于衡量模型在实际协作中的表现
大型语言模型正快速融入开发闭环,从辅助工具演变为深度嵌入协作环境的自主贡献者。然而现有评估方法因碎片化和对真实能力的扭曲呈现而受限,常基于假设性语法场景。本研究旨在填补这一空白,提出一种基于真实软件开发实践的综合评估方法。评估体系聚焦污染感知、野外环境下代理行为评估,以及捕捉真实编码情境、人类对齐行为和模型失效模式的轨迹感知基准与度量指标。
原文摘要 · Abstract (English)
Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive evaluation methodology for LLM-powered agents that is grounded in real-world software development practice. Our evaluation approach focuses on contamination-awareness, in-the-wild agentic behavior assessment, and trajectory-aware benchmarks and metrics capturing realistic coding contexts, human-aligned behavior, and model failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。