arXiv:2412.03775cs.CLcs.DL2024-12被引 12

首个大规模撤稿论文数据集,助力科研诚信研究。

WithdrarXiv: A Large-Scale Dataset for Retraction Study

  • 构建包含1.4万篇论文的arXiv撤稿数据集,涵盖完整历史。
  • 自动分类准确率达F1=0.96,可识别10类撤稿原因。
  • 提供解析版全文脚本,支持科学可行性验证研究。

撤稿在维护科学诚信中至关重要,但计算机科学及其他STEM领域系统的撤稿研究仍很匮乏。本文提出WithdrarXiv,首个来自arXiv的大规模撤稿论文数据集,包含截至2024年9月超过14,000篇论文及其相关的撤稿评论。通过分析作者说明,我们建立了一个全面的撤稿原因分类体系,识别出10类不同原因,从重大错误到政策违规。我们展示了一种简单但高度准确的零样本自动分类方法,达到加权平均F1分数0.96。此外,我们发布WithdrarXiv-SciFy版本,包含已解析的全文PDF脚本,专为科学可行性研究、主张验证和自动化定理证明设计。这些成果为提升科学质量控制与自动化验证系统提供了重要洞见。最后,我们讨论了伦理问题,并采取多项措施实现负责任的数据发布,推动该领域的开放科学。

原文摘要 · Abstract (English)

Retractions play a vital role in maintaining scientific integrity, yet systematic studies of retractions in computer science and other STEM fields remain scarce. We present WithdrarXiv, the first large-scale dataset of withdrawn papers from arXiv, containing over 14,000 papers and their associated retraction comments spanning the repository's entire history through September 2024. Through careful analysis of author comments, we develop a comprehensive taxonomy of retraction reasons, identifying 10 distinct categories ranging from critical errors to policy violations. We demonstrate a simple yet highly accurate zero-shot automatic categorization of retraction reasons, achieving a weighted average F1-score of 0.96. Additionally, we release WithdrarXiv-SciFy, an enriched version including scripts for parsed full-text PDFs, specifically designed to enable research in scientific feasibility studies, claim verification, and automated theorem proving. These findings provide valuable insights for improving scientific quality control and automated verification systems. Finally, and most importantly, we discuss ethical issues and take a number of steps to implement responsible data release while fostering open science in this area.

数据集科研诚信自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。