arXiv:2506.11343cs.CL2025-06NeurIPS综述被引 13

用LLM对比论文而非打分,更准识别高影响力文章

From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review

  • 让LLM代理两两比较论文,通过投票决定优劣
  • 相比传统打分法,识别高影响力论文准确率显著提升
  • 适合关注学术评审革新与AI公平性的研究者

大型语言模型(LLMs)为重构学术同行评审提供了前所未有的机遇。然而,现有工作多局限于将LLM作为人类审稿人的直接替代,缺乏对新范式的探索。本文提出并验证一种新机制:让LLM代理对论文进行两两比较,而非单独评分。通过聚合大量比较结果,可获得更准确、更稳健的相对质量评估。实验表明,该方法在识别高影响力论文方面显著优于传统评分方式。但分析也揭示了潜在偏见,如研究主题新颖性下降、机构间不平衡加剧。这些发现凸显了以LLM重构评审的变革潜力,也指出了未来系统需应对的公平性与多样性挑战。

原文摘要 · Abstract (English)

The advent of large language models (LLMs) offers unprecedented opportunities to reimagine peer review beyond the constraints of traditional workflows. Despite these opportunities, prior efforts have largely focused on replicating traditional review workflows with LLMs serving as direct substitutes for human reviewers, while limited attention has been given to exploring new paradigms that fundamentally rethink how LLMs can participate in the academic review process. In this paper, we introduce and explore a novel mechanism that employs LLM agents to perform pairwise comparisons among manuscripts instead of individual scoring. By aggregating outcomes from substantial pairwise evaluations, this approach enables a more accurate and robust measure of relative manuscript quality. Our experiments demonstrate that this comparative approach significantly outperforms traditional rating-based methods in identifying high-impact papers. However, our analysis also reveals emergent biases in the selection process, notably a reduced novelty in research topics and an increased institutional imbalance. These findings highlight both the transformative potential of rethinking peer review with LLMs and critical challenges that future systems must address to ensure equity and diversity.

LLM同行评审对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。