arXiv:2601.15172cs.CL2026-01综述

用新方法分析三大顶会审稿质量,发现没明显下降

Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time

  • 构建多维度评分框架,结合大模型与轻量指标量化审稿价值
  • 跨年对比显示主流会议审稿质量无持续下滑趋势
  • 适合关注学术评审机制的科研人员和期刊编辑参考

同行评审是现代科学的核心。随着投稿量增加和研究社区扩张,审稿质量下降成为普遍叙事和常见担忧。然而,这一说法是否成立?审稿质量难以衡量,且评审实践持续演变,导致跨会议、跨时间比较困难。为此,我们提出一种基于证据的审稿质量比较新框架,并应用于人工智能与机器学习领域的三大顶会:ICLR、NeurIPS 和 ACL。我们记录了审稿格式的多样性,提出一种新的审稿标准化方法。设计多维度评分体系,以评估审稿对编辑和作者的实用性,并采用基于大语言模型(LLM)和轻量级方法的双重测量。研究分析了不同度量方式之间的关系及其随时间的变化。与流行观点相反,我们的跨时间分析未发现主要会议在多年间中位审稿质量存在一致下降。我们提出替代解释,并为未来实证研究审稿质量提供改进建议。

原文摘要 · Abstract (English)

Peer review is at the heart of modern science. As submission numbers rise and research communities grow, the decline in review quality is a popular narrative and a common concern. Yet, is it true? Review quality is difficult to measure, and the ongoing evolution of reviewing practices makes it hard to compare reviews across venues and time. To address this, we introduce a new framework for evidence-based comparative study of review quality and apply it to major AI and machine learning conferences: ICLR, NeurIPS and *ACL. We document the diversity of review formats and introduce a new approach to review standardization. We propose a multi-dimensional schema for quantifying review quality as utility to editors and authors, coupled with both LLM-based and lightweight measurements. We study the relationships between measurements of review quality, and its evolution over time. Contradicting the popular narrative, our cross-temporal analysis reveals no consistent decline in median review quality across venues and years. We propose alternative explanations, and outline recommendations to facilitate future empirical studies of review quality.

同行评审量化评估顶会分析学术诚信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。