arXiv:2607.27209cs.DLcs.AI2026-07综述被引 6

ML论文评审分数跨领域不可比,导致录用率差异高达8倍。

Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

论文配图:Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
图 1 · 摘自论文原文
  • 用统一分数尺度过滤不同研究领域,造成评分失真。
  • 同一分数下,不同领域的录用概率最高差8倍。
  • 建议公开分主题的录用率数据,提升评审公平性。

机器学习会议的同行评审日益依赖评审分数作为核心决策依据。随着投稿量从数千激增至数万篇,尚无系统性审计检验该工具在不同研究领域是否一致有效,或录用结果是否受评分无法捕捉与控制的因素影响。本文指出,录用结果受评分之外因素主导,根源在于测量设计缺陷:当固定数值尺度用于聚合结构差异大的评审群体时,绝对分数在不同领域失去可比性,领域主席不得不以社区先验替代分数决策。基于ICLR 2021–2026年覆盖219个研究主题、共50,289篇论文的数据,我们发现,在相同评审分数下,论文录用概率因主题不同最高相差8倍。我们排除了评分文化、专家标准、理性重加权及质量稀释等替代解释。呼吁程序委员会采用内在校准的评审信号,并公布按主题分层的条件录用率,将其作为首要公平性指标。

原文摘要 · Abstract (English)

Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021--2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper's acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.

同行评审公平性机器学习量化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。