arXiv:2602.00733cs.CLcs.AI2026-02综述被引 1

利用论文引用上下文构建自动审稿数据,提升审稿质量与可扩展性。

EchoReview: Learning Peer Review from the Echoes of Scientific Citations

  • 通过分析学术引用上下文,挖掘科学共同体长期评价信号作为审稿数据。
  • 构建首个跨会议跨年度的16K规模引用驱动审稿数据集,训练出70亿参数模型。
  • 模型在证据支持和审稿全面性上显著优于传统方法,适合大规模科研评审场景。

随着科学投稿量快速增长,传统同行评审系统面临前所未有的可扩展性压力,亟需既高效又可靠的自动化评审方法。现有基于真实评审数据的监督微调方法受限于单一数据源及人工评审的主观性和不一致性,难以支撑高质量自动化评审。为此,我们提出EchoReview,一种基于引用上下文的数据合成框架,系统性地挖掘学术引用中隐含的集体评价信号,并将科学共同体长期判断转化为结构化审稿数据。基于此流程,我们构建了首个大规模、跨会议、跨年的引用驱动审稿数据集EchoReview-16K,训练出自动化评审模型EchoReviewer-7B。实验表明,EchoReviewer-7B在证据支持、审稿全面性等核心维度上实现显著且稳定的提升,验证了引用上下文作为可靠自动化同行评审数据范式的有效性。

原文摘要 · Abstract (English)

As the volume of scientific submissions continues to grow rapidly, traditional peer review systems are facing unprecedented scalability pressures, highlighting the urgent need for automated reviewing methods that are both scalable and reliable. Existing supervised fine-tuning approaches based on real review data are fundamentally constrained by single-source of data as well as the inherent subjectivity and inconsistency of human reviews, limiting their ability to support high-quality automated reviewers. To address these issues, we propose EchoReview, a citation-context-driven data synthesis framework that systematically mines implicit collective evaluative signals from academic citations and transforms scientific community's long-term judgments into structured review-style data. Based on this pipeline, we construct EchoReview-16K, the first large-scale, cross-conference, and cross-year citation-driven review dataset, and train an automated reviewer, EchoReviewer-7B. Experimental results demonstrate that EchoReviewer-7B can achieve significant and stable improvements on core review dimensions such as evidence support and review comprehensiveness, validating citation context as a robust and effective data paradigm for reliable automated peer review.

同行评审引用分析自动化数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。