arXiv:2502.00902cs.SEcs.LG2025-02被引 1

ML研究复现难,根源在于软件工程不规范。

More Rigorous Software Engineering Would Improve Reproducibility in Machine Learning Research

  • 调研近十年顶会论文代码库,发现软件实践普遍缺失
  • 仅12%代码库使用版本控制,37%缺乏自动化测试
  • 建议推广标准化开发流程,提升研究可信度

尽管实验复现是科学方法的核心,我们观察到支持机器学习(ML)研究复现的软件工程最佳实践常被忽视,导致复现性差并损害了ML社区的信任。通过调查过去十年在NeurIPS、ICML、ICLR、TMLR和MLOSS等主要会议和期刊发表论文所关联的代码仓库,我们量化了软件最佳实践的使用情况。结果显示,仅有12%的仓库使用版本控制,37%缺乏自动化测试,且文档和依赖管理普遍不足。本文识别出当前存在的薄弱环节,并提出具体改进措施,呼吁学术界共同建立更严谨的软件开发规范,以提升机器学习研究的可复现性。

原文摘要 · Abstract (English)

While experimental reproduction remains a pillar of the scientific method, we observe that the software best practices supporting the reproduction of machine learning ( ML ) research are often undervalued or overlooked, leading both to poor reproducibility and damage to trust in the ML community. We quantify these concerns by surveying the usage of software best practices in software repositories associated with publications at major ML conferences and journals such as NeurIPS, ICML, ICLR, TMLR, and MLOSS within the last decade. We report the results of this survey that identify areas where software best practices are lacking and areas with potential for growth in the ML community. Finally, we discuss the implications and present concrete recommendations on how we, as a community, can improve reproducibility in ML research.

可复现性软件工程研究规范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。