首次大规模研究大模型在软件工程中的可复现性问题
Large Language Models for Software Engineering: A Reproducibility Crisis
- 系统分析640篇论文的可复现性缺陷,发现代码、环境、文档等多方面缺失
- 尽管近年有改善,但80%以上论文仍存在可复现性问题,认证徽章不等于真正可运行
- 提出可复现性成熟度模型,推动从“有无”到“好坏”的评估进步
可复现性是科学进步的基础,但大语言模型用于软件工程(LLM-for-SE)研究中的可复现性状况尚不明确。本文首次开展大规模实证研究,系统挖掘并分析了2017至2025年间在顶级软件工程、机器学习与自然语言处理会议发表的640篇论文,提取出版物、代码仓库与文档中的结构化元数据。基于四个研究问题:(i)可复现性缺陷的普遍性;(ii)可复现性随时间的变化;(iii)成果评价徽章是否真实反映可复现质量;(iv)发表会议对透明度的影响。我们采用包含七类缺陷的分类体系——代码与执行、数据、文档、环境与工具、版本管理、模型、访问与法律——对所有论文及关联成果进行人工标注。分析发现,尽管近年有所改善且顶级会议更多采纳成果评价流程,但在代码可用性、环境说明、版本严谨性和文档清晰性方面仍存在持续缺口。值得注意的是,徽章仅表明成果存在,却不能保证执行一致或长期可复现。基于此,我们提出可操作建议,并引入可复现性成熟度模型(RMM),推动从二元认证转向多维度、渐进式的可复现性评估。
原文摘要 · Abstract (English)
Reproducibility is a cornerstone of scientific progress, yet its state in large language model (LLM)-based software engineering (SE) research remains poorly understood. This paper presents the first large-scale, empirical study of reproducibility practices in LLM-for-SE research. We systematically mined and analyzed 640 papers published between 2017 and 2025 across premier software engineering, machine learning, and natural language processing venues, extracting structured metadata from publications, repositories, and documentation. Guided by four research questions, we examine (i) the prevalence of reproducibility smells, (ii) how reproducibility has evolved over time, (iii) whether artifact evaluation badges reliably reflect reproducibility quality, and (iv) how publication venues influence transparency practices. Using a taxonomy of seven smell categories: Code and Execution, Data, Documentation, Environment and Tooling, Versioning, Model, and Access and Legal, we manually annotated all papers and associated artifacts. Our analysis reveals persistent gaps in artifact availability, environment specification, versioning rigor, and documentation clarity, despite modest improvements in recent years and increased adoption of artifact evaluation processes at top SE venues. Notably, we find that badges often signal artifact presence but do not consistently guarantee execution fidelity or long-term reproducibility. Motivated by these findings, we provide actionable recommendations to mitigate reproducibility smells and introduce a Reproducibility Maturity Model (RMM) to move beyond binary artifact certification toward multi-dimensional, progressive evaluation of reproducibility rigor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。