让AI学会像专家一样反复审稿,提前预测论文修改周期。
FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes

- 用多轮审稿对话训练模型,模拟真实科学评审的迭代过程。
- 准确预测80.5%的审稿结果,显著优于现有模型10.4个百分点。
- 适合科研作者在投稿前自测,提升论文质量与通过率。
当前AI同行评审系统存在三大缺陷:仅基于计算机科学与机器学习会议训练、忽略科学验证中的互动对话、评估标准偏向形式模仿而非真实编辑判断。本文提出FirstPass,包含3,668条来自《自然·通讯》五个学科(生物学、化学、神经科学、物理学、地球科学)的完整多轮审稿对话数据集,并利用2022年11月起强制透明审稿机制确保内容完整性。采用低秩适配(LoRA)微调Qwen2.5-7B-Instruct模型,完成三项任务:审稿生成、审稿人更新与修订周期预测。关键发现是‘仅响应损失掩码’为必要前提而非优化:无此设计时准确率仅62.0%,低于多数基线;有此设计时,对审稿结果预测达到80.5%准确率和78.2% F1-macro,较Gemini-3.1-flash-lite-preview零样本高出10.4个百分点,且统计显著(McNemar p < 0.001)。在生成任务中,FirstPass平均生成1,187字审稿意见,接近人类参考(2,155字),ROUGE-L达0.154,显著优于Qwen与DeepSeek零样本(p < 0.001)。作为预提交阶段的前瞻性科学合作者,FirstPass可模拟专家批判并预测修订周期,跨五领域表现稳定。
原文摘要 · Abstract (English)
AI systems for peer review fail on three fronts: they train on Computer Science and Machine Learning venues alone, ignore the iterative dialogue that validates science, and evaluate on stylistic mimicry rather than real editorial judgment. We introduce FirstPass, a dataset and fine-tuned model that addresses all three. Curating 3,668 complete multi-round peer-review dialogues from Nature Communications across five scientific domains (biology, chemistry, neuroscience, physics, and earth science), we exploit mandatory transparent peer review (instituted November 2022) and verify 100% content integrity by automated audit. We fine-tune Qwen2.5-7B-Instruct via Low-Rank Adaptation (LoRA) on three tasks: review generation, reviewer updating, and revision-cycle prediction. Our key finding is that response-only loss masking is a prerequisite, not an optimization: without it, accuracy is 62.0%, below the majority baseline; with it, FirstPass achieves 80.5% accuracy and F1-macro 78.2% on predicting editorial outcomes (Standard vs. Extended revision cycles), outperforming Gemini-3.1-flash-lite-preview zero-shot by 10.4 percentage points and all baselines with statistical significance (McNemar p < 0.001). On generation, FirstPass produces reviews averaging 1,187 words, substantially closer to human references (2,155 words) than any baseline, achieving ROUGE-L 0.154 with significant gains over Qwen and DeepSeek zero-shot (p < 0.001). Deployed in the pre-submission loop as an anticipatory scientific co-author, FirstPass simulates expert critique and predicts revision cycle outcomes before submission, giving authors the judgment a trusted colleague would provide, with consistent cross-domain performance across five disciplines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。