arXiv:2510.25506cs.SEcs.AI2025-10被引 14

85篇顶会论文仅5篇可复现,揭示大模型研究可复现性危机

Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies

  • 分析85篇顶会大模型研究,仅18篇提供完整代码和数据
  • 18篇中仅5篇可执行,全都不完全复现原始结果
  • 呼吁强化代码公开与实验设计,提升科研可信度

大型语言模型在产业界和学术界均引发广泛关注。以ICSE 2024为例,约425篇论文中就有78篇涉及大模型实验。尽管研究热度上升,但实证研究仍面临可复现性挑战。本文分析了发表于ICSE 2024和ASE 2024的85篇大模型相关研究,其中18篇提供了研究数据和使用OpenAI模型。我们尝试复现这18项研究,发现仅5篇具备可执行性。然而,对这5项研究均未能完全复现原始结果;另有两项部分可复现,三项不可复现。研究揭示当前大模型研究存在严重可复现性问题,强调需加强研究数据与代码审查,并改进实验设计以确保未来成果的可靠性。

原文摘要 · Abstract (English)

Large Language Models have gained remarkable interest in industry and academia. The increasing interest in LLMs in academia is also reflected in the number of publications on this topic over the last years. For instance, alone 78 of the around 425 publications at ICSE 2024 performed experiments with LLMs. Conducting empirical studies with LLMs remains challenging and raises questions on how to achieve reproducible results, for both researchers and practitioners. One important step towards excelling in empirical research on LLM and their application is to first understand to what extent current research results are eventually reproducible and what factors may impede reproducibility. This investigation is within the scope of our work. We contribute an analysis of the reproducibility of LLM-centric studies, provide insights into the factors impeding reproducibility, and discuss suggestions on how to improve the current state. In particular, we studied the 85 articles describing LLM-centric studies, published at ICSE 2024 and ASE 2024. Of the 85 articles, 18 provided research artefacts and used OpenAI models. We attempted to replicate those 18 studies. Of the 18 studies, only five were sufficiently complete and executable. For none of the five studies, we were able to fully reproduce the results. Two studies seemed to be partially reproducible, and three studies did not seem to be reproducible. Our results highlight not only the need for stricter research artefact evaluations but also for more robust study designs to ensure the reproducible value of future publications.

可复现性大模型软件工程实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。