50%的顶尖AI论文可复现,开源数据最关键。
The Unreasonable Effectiveness of Open Science in AI: A Replication Study
- 系统复现30篇高被引AI论文,仅8篇因数据/硬件不可得被拒。
- 共享代码与数据的论文复现率达86%,远超仅共享数据的33%。
- 数据文档质量决定成败,代码文档好坏无关复现成功与否。
人工智能研究是否存在可复现性危机尚不明确。为此,我们对30篇高被引AI论文进行了系统性复现研究,尽可能使用原始材料。最终,8篇因无法获取数据或硬件而被拒绝;6篇完全复现,5篇部分复现,总计50%的论文得以复现。代码与数据共享显著提升复现率:86%共享代码与数据的论文被完全或部分复现,而仅共享数据的论文复现率仅为33%。数据文档质量与复现成功率强相关,文档不清或数据定义错误将导致复现失败。令人意外的是,代码文档质量与复现结果无关——无论代码文档差、缺失或未版本化,只要代码公开,复现仍可能成功。本研究凸显开放科学的有效性及数据工作规范记录的重要性。
原文摘要 · Abstract (English)
A reproducibility crisis has been reported in science, but the extent to which it affects AI research is not yet fully understood. Therefore, we performed a systematic replication study including 30 highly cited AI studies relying on original materials when available. In the end, eight articles were rejected because they required access to data or hardware that was practically impossible to acquire as part of the project. Six articles were successfully reproduced, while five were partially reproduced. In total, 50% of the articles included was reproduced to some extent. The availability of code and data correlate strongly with reproducibility, as 86% of articles that shared code and data were fully or partly reproduced, while this was true for 33% of articles that shared only data. The quality of the data documentation correlates with successful replication. Poorly documented or miss-specified data will probably result in unsuccessful replication. Surprisingly, the quality of the code documentation does not correlate with successful replication. Whether the code is poorly documented, partially missing, or not versioned is not important for successful replication, as long as the code is shared. This study emphasizes the effectiveness of open science and the importance of properly documenting data work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。