为自动化复现难题构建通用问题表述,让机器能理解研究核心。
Automated Reproducibility Has a Problem Statement Problem
- 用科学方法结构化表示20项实证研究,自动提取假设、实验和结论。
- 原作者对90%以上结构表示认可,但细节如实验结果仍需改进。
- 适合关注可复现性、AI研究自动化与论文结构化的研究者。
可复现性是科学方法的核心,但复现实验常耗时费力。近期研究尝试自动化该过程以减轻负担,但因复现定义不一,缺乏清晰的问题表述。本文旨在为任意实证研究创建可泛化的复现问题表述。我们假设:任何实证研究均可基于科学方法建模,并能自动从文献中提取其假设、实验与解释。通过自动提取20项研究的上述要素,并由原作者评估质量,我们构建了一个包含20项研究的复现问题数据集。多数作者对提取结构表示认可,涵盖各部分;少数案例未能完整捕捉研究内容,尤其在实验结果细节上仍有提升空间。结论表明,该问题表述能有效覆盖广泛子领域的实证人工智能研究。原作者普遍认为生成结构具有代表性,未来可通过更精细输出进一步优化。
原文摘要 · Abstract (English)
Background. Reproducibility is essential to the scientific method, but reproduction is often a laborious task. Recent works have attempted to automate this process and relieve researchers of this workload. However, due to varying definitions of reproducibility, a clear problem statement is missing. Objectives. Create a generalisable problem statement, applicable to any empirical study. We hypothesise that we can represent any empirical study using a structure based on the scientific method and that this representation can be automatically extracted from any publication, and captures the essence of the study. Methods. We apply our definition of reproducibility as a problem statement for the automatisation of reproducibility by automatically extracting the hypotheses, experiments and interpretations of 20 studies and assess the quality based on assessments by the original authors of each study. Results. We create a dataset representing the reproducibility problem, consisting of the representation of 20 studies. The majority of author feedback is positive, for all parts of the representation. In a few cases, our method failed to capture all elements of the study. We also find room for improvement at capturing specific details, such as results of experiments. Conclusions. We conclude that our formulation of the problem is able to capture the concept of reproducibility in empirical AI studies across a wide range of subfields. Authors of original publications generally agree that the produced structure is representative of their work; we believe improvements can be achieved by applying our findings to create a more structured and fine-grained output in future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。