arXiv:2605.26730cs.CL2026-05被引 3

PRISM benchmarks LLM评审器在四大维度的表现,发现其各有专长但无法全面替代人类。

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

论文配图:PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
图 1 · 摘自论文原文
  • 分四个维度评估评审质量,基于论证挖掘与共识评分。
  • LLM在新颖性识别和问题优先级上优于人类,但整体平衡性不足。
  • 适合研究者评估自动化评审系统,不建议独立使用。

机器学习会议投稿量激增,促使学界关注基于大模型的自动化评审系统。然而,这些系统的真实表现,尤其是与人类相比能否发现科学漏洞,仍不清晰。本文提出PRISM(Peer Review Intelligence via Structured Multi-dimensional assessment),一个涵盖深度分析、新颖性评估、缺陷识别与重大问题优先级、多维建设性四维度的评估框架。不同于依赖ROUGE、BLEU等表面指标或自由式LLM评判的方法,PRISM通过论证挖掘、检索增强验证和共识评分实现严谨评估。我们在ICLR、ICML、NeurIPS的分层评审语料上测试了五个主流自动评审系统及人类评审员。结果表明:LLM在单一维度上可媲美甚至超越人类——深度分析相当,新颖性验证更强,批判性问题优先级准确;但无一系统能在所有维度上持续达到人类评审的整体平衡表现。各系统均有特定专长与盲区,聚合指标难以捕捉此类失效模式。结论是,LLM评审器应视为人类评审的靶向补充,而非独立替代品。演示与关键结果见https://khanhthanhdev.github.io/prism-page/。

原文摘要 · Abstract (English)

The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automated peer reviewers. However, how good these systems are actually, especially compared to human reviewers at catching scientific gaps, remains poorly understood. In this work, we introduce PRISM (Peer Review Intelligence via Structured Multi-dimensional assessment), a benchmarking framework that evaluates review quality across four dimensions: Depth of Analysis, Novelty Assessment,Flaw Identification & Major Issues Prioritization, and Multi-dimensional Constructiveness. Unlike most existing evaluations based on surface-level metrics like ROUGE and BLEU, or unconstrained LLM-as-a-judge prompting that conflates fluency with rigor, PRISM grounds each dimension in argument mining, retrieval-augmented verification, and consensus-based scoring. We apply PRISM to benchmark five leading automated reviewer systems and human reviewers on a stratified corpus of reviews from ICLR, ICML, and NeurIPS. The results reveal that LLMs can match or beat human reviewers on individual dimensions: comparable depth of analysis, stronger novelty verification, and highly accurate critique prioritization. However, no single system consistently matches the balanced performance of the human baseline across all dimensions at once. Each exhibits a distinct specialization profile with characteristic blind spots -- failure modes that aggregate metrics miss entirely. The implication is that LLM reviewers are best understood as targeted supplements to human review, effective within specific dimensions, but unreliable as standalone replacements. Our demo and key results can be found at https://khanhthanhdev.github.io/prism-page/.

大模型评审评测基准AI辅助学术出版

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。