arXiv:2507.17991cs.SEcs.IR2025-07

对比11个工具检测科研透明度,发现组合使用效果更好。

Use as Directed? A Comparison of Software Tools Intended to Check Rigor and Transparency of Published Work

  • 用11个自动化工具评估9项科研严谨性标准
  • 组合工具在开放数据检测上表现显著优于单个工具
  • 适合期刊、审稿人和工具开发者参考优化

科研可重复性危机的根源之一是报告缺乏标准化与透明度。虽然ARRIVE和CONSORT等清单旨在提升透明度,但作者常未遵循,同行评审也难以发现缺失项。为此,已有多个自动化工具被设计用于检测不同严谨性指标。本研究对11个工具在ScreenIT提出的9项严谨性标准上进行了全面比较。结果显示,在开放数据检测等部分标准上,某些工具明显优于其他工具;而在纳入/排除标准检测等任务中,组合使用多个工具的效果超越单一工具。研究还识别出工具开发应重点改进的方向,并为相关利益方提供了优化建议。研究代码与数据已公开于https://github.com/PeterEckmann1/tool-comparison。

原文摘要 · Abstract (English)

The causes of the reproducibility crisis include lack of standardization and transparency in scientific reporting. Checklists such as ARRIVE and CONSORT seek to improve transparency, but they are not always followed by authors and peer review often fails to identify missing items. To address these issues, there are several automated tools that have been designed to check different rigor criteria. We have conducted a broad comparison of 11 automated tools across 9 different rigor criteria from the ScreenIT group. We found some criteria, including detecting open data, where the combination of tools showed a clear winner, a tool which performed much better than other tools. In other cases, including detection of inclusion and exclusion criteria, the combination of tools exceeded the performance of any one tool. We also identified key areas where tool developers should focus their effort to make their tool maximally useful. We conclude with a set of insights and recommendations for stakeholders in the development of rigor and transparency detection tools. The code and data for the study is available at https://github.com/PeterEckmann1/tool-comparison.

科研透明度自动化检测工具对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。