arXiv:2511.07790cs.DLcs.CL2025-11中稿 · JCDL 2025, 16 page…

构建3万条机器学习论文引用语境数据集,用于评估研究可复现性情绪。

CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis

  • 通过众包与可控生成结合,构建含30,734条标注的可复现性情感数据集。
  • 标签准确率达94%,三款大模型在微调后分类性能显著提升。
  • 适合关注科研可信度、可复现性评估的研究者使用。

关于被引论文可复现性的观点能反映学术社区的看法,并已被证明是实际研究成果可复现性的潜在信号。为训练有效模型以预测可复现性导向的情感,并进一步系统研究其与可复现性的关联,我们引入了CC30k数据集,包含30,734条机器学习论文中的引用语境。每条引用语境标注为三种可复现性导向情感之一:正面、负面或中性,反映对被引论文可复现性或可重复性的感知。其中25,829条通过众包标注,其余负样本通过受控流水线生成以缓解负样本稀缺问题。与传统情感分析数据集不同,CC30k聚焦于可复现性导向情感,填补了计算可复现性研究资源的空白。数据集通过包含强数据清洗、精心挑选众包人员和全面验证的流程构建,标签准确率达94%。我们还证明,三款大语言模型在使用本数据集微调后,在可复现性导向情感分类任务上性能显著提升。该数据集为大规模评估机器学习论文可复现性奠定了基础。CC30k数据集及用于生成与分析数据集的Jupyter笔记本已公开发布于https://github.com/lamps-lab/CC30k。

原文摘要 · Abstract (English)

Sentiments about the reproducibility of cited papers in downstream literature offer community perspectives and have shown as a promising signal of the actual reproducibility of published findings. To train effective models to effectively predict reproducibility-oriented sentiments and further systematically study their correlation with reproducibility, we introduce the CC30k dataset, comprising a total of 30,734 citation contexts in machine learning papers. Each citation context is labeled with one of three reproducibility-oriented sentiment labels: Positive, Negative, or Neutral, reflecting the cited paper's perceived reproducibility or replicability. Of these, 25,829 are labeled through crowdsourcing, supplemented with negatives generated through a controlled pipeline to counter the scarcity of negative labels. Unlike traditional sentiment analysis datasets, CC30k focuses on reproducibility-oriented sentiments, addressing a research gap in resources for computational reproducibility studies. The dataset was created through a pipeline that includes robust data cleansing, careful crowd selection, and thorough validation. The resulting dataset achieves a labeling accuracy of 94%. We then demonstrated that the performance of three large language models significantly improves on the reproducibility-oriented sentiment classification after fine-tuning using our dataset. The dataset lays the foundation for large-scale assessments of the reproducibility of machine learning papers. The CC30k dataset and the Jupyter notebooks used to produce and analyze the dataset are publicly available at https://github.com/lamps-lab/CC30k .

可复现性情感分析数据集机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。