arXiv:2608.26129cs.CLcs.AI2026-08中稿 · the AI for Science…综述

首个跨学科多轮审稿数据集,真实还原期刊编辑决策过程。

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

论文配图:FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
图 1 · 摘自论文原文
  • 基于《自然·通讯》透明审稿记录构建,覆盖五大学科
  • 含3668条完整审稿对话,每条带编辑决策标签
  • 适合研究AI科学判断、跨学科评审差异的学者使用

现有科学审稿数据集仅涵盖计算机与机器学习领域,导致模型只能评估消融实验,却从未见过生物学家要求污染控制或化学家质疑核磁共振谱图归属。我们提出FIRSTPASS,首个基于多轮编辑对话的大规模跨学科审稿数据集,源自《自然·通讯》自2022年11月起强制推行的透明审稿机制。数据涵盖生物学、化学、神经科学、物理学和地球科学五个领域,共3668条记录,完整呈现科学验证的迭代过程:初审意见、作者逐点回应及更新评审。每条记录均标注来自编辑决策的真实结果(两轮为STANDARD,三轮及以上为EXTENDED),提供此前所有数据集缺失的金标准。自动化审计确认内容完整性达100%。专家审稿平均字数达2155字,显著高于会议审稿。所有数据、解析流程与评估脚本均已公开,支持跨学科人工智能科学判断的可复现基准测试。

原文摘要 · Abstract (English)

Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal. Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments. Each record carries an outcome label derived directly from editorial decisions (STANDARD for two-round review; EXTENDED for three or more rounds), providing ground truth absent in all prior corpora. An automated audit confirms 100% content integrity. Expert reviews average 2,155 words, substantially denser than conference venue reviews. All data, parsing pipelines, and evaluation scripts are released to enable reproducible benchmarking of AI scientific judgment across disciplines.

审稿数据集跨学科AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。