用开放评审数据量化审稿人质量,提升论文评估的准确性和公平性。
Paper Quality Assessment based on Individual Wisdom Metrics from Open Peer Review
- 基于审稿人评分与社区共识的差异,构建个体审稿质量评估指标。
- 发现中等水平论文作者反而是最可靠的审稿人,呈现倒U型关系。
- 采用贝叶斯方法融合可信度,显著优于简单平均,适合大规模开放评审。
传统封闭式同行评审存在速度慢、透明度低、评分波动大及潜在偏见等问题,阻碍科学进步并削弱公众信任。本文基于CCN2023和ICLR2023两个大会的数据,提出一种自下而上的开放评审机制。研究发现,审稿人评分差异大且相关性低,难以形成有效集体判断。通过量化审稿人与社区共识的一致性,构建其质量评估指标,结果显示审稿人质量与作者论文质量无相关性,反而中等论文作者的审稿表现最佳,呈倒U型分布。采用经验贝叶斯方法融合不同审稿人可靠性,证明在单次评审场景下,该方法显著优于简单平均。进一步考虑持续评分机制(审稿人同时评价其他审稿人),即使多数审稿人不可靠但无偏,用户生成的评审评分仍能保持高质量。最后提出激励机制以表彰优质审稿人,促进更广泛覆盖投稿论文。结果表明,自选参与的开放评审具备可扩展性、可靠性与公平性,有望提升评审效率、公正性与透明度。
原文摘要 · Abstract (English)
Traditional closed peer review systems, which have played a central role in scientific publishing, are often slow, costly, non-transparent, stochastic, and possibly subject to biases - factors that can impede scientific progress and undermine public trust. Here, we propose and examine the efficacy and accuracy of an alternative form of scientific peer review: through an open, bottom-up process. First, using data from two major scientific conferences (CCN2023 and ICLR2023), we highlight how high variability of review scores and low correlation across reviewers presents a challenge for collective review. We quantify reviewer agreement with community consensus scores and use this as a reviewer quality estimator, showing that surprisingly, reviewer quality scores are not correlated with authorship quality. Instead, we reveal an inverted U-shape relationship, where authors with intermediate paper scores are the best reviewers. We assess empirical Bayesian methods to estimate paper quality based on different assessments of individual reviewer reliability. We show how under a one-shot review-then-score scenario, both in our models and on real peer review data, a Bayesian measure significantly improves paper quality assessments relative to simple averaging. We then consider an ongoing model of publishing, reviewing, and scoring, with reviewers scoring not only papers but also other reviewers. We show that user-generated reviewer ratings can yield robust and high-quality paper scoring even when unreliable (but unbiased) reviewers dominate. Finally, we outline incentive structures to recognize high-quality reviewers and encourage broader reviewing coverage of submitted papers. These findings suggest that a self-selecting open peer review process is potentially scalable, reliable, and equitable with the possibility of enhancing the speed, fairness, and transparency of the peer review process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。