提出可复现可解释的智能体评估指南,提升软件工程中AI研究可信度。
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
- 建议公开思维-动作-结果轨迹和LLM交互数据
- 分析18篇顶会论文发现评估存在不可复现问题
- 适合关注AI可解释性与实验严谨性的研究者
随着智能体AI的发展,研究人员越来越多地使用自主智能体解决软件工程(SE)挑战。然而,支撑这些智能体的大语言模型(LLMs)常被视为黑箱,难以证明智能体方法优于基线。此外,评估设计描述不完整,常导致结果无法复现。本研究分析了18篇发表于ICSE 2026、ICSE 2025、FSE 2025、ASE 2025和ISSTA 2025的论文,梳理当前智能体AI在软件工程中的评估实践及其局限。为改善这些问题,本文提出一套指南,旨在推动可复现、可解释、有效的智能体AI评估。特别建议研究者公开其思维-动作-结果(TAR)轨迹及LLM交互数据或其摘要版本,以支持后续系统性分析。通过一个概念验证案例,展示了如何利用TAR轨迹进行跨方法比较。
原文摘要 · Abstract (English)
With the advancement of Agentic AI, researchers are increasingly leveraging autonomous agents to address challenges in software engineering (SE). However, the large language models (LLMs) that underpin these agents often function as black boxes, making it difficult to justify the superiority of Agentic AI approaches over baselines. Furthermore, missing information in the evaluation design description frequently renders the reproduction of results infeasible. To synthesize current evaluation practices for Agentic AI in SE, this study analyzes 18 papers on the topic, published or accepted by ICSE 2026, ICSE 2025, FSE 2025, ASE 2025, and ISSTA 2025. The analysis identifies prevailing approaches and their limitations in evaluating Agentic AI for SE, both in current research and potential future studies. To address these shortcomings, this position paper proposes a set of guidelines and recommendations designed to empower reproducible, explainable, and effective evaluations of Agentic AI in software engineering. In particular, we recommend that Agentic AI researchers make their Thought-Action-Result (TAR) trajectories and LLM interaction data, or summarized versions of these artifacts, publicly accessible. Doing so will enable subsequent studies to more effectively analyze the strengths and weaknesses of different Agentic AI approaches. To demonstrate the feasibility of such comparisons, we present a proof-of-concept case study that illustrates how TAR trajectories can support systematic analysis across approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。