arXiv:2608.03659cs.CLcs.AI2026-08综述

对比三款大模型与人工审稿,发现评分不一致且关注点不同。

How Closely Do LLM Reviews Align with Human Peer Review?

论文配图:How Closely Do LLM Reviews Align with Human Peer Review?
图 1 · 摘自论文原文
  • 用统一指令让三款大模型审300篇ICLR 2026论文,去决策信息后比对
  • 模型能区分接受与拒绝,但无法复现人工对口头报告与海报的区分
  • 模型更关注基线缺失,人工更在意计算效率,评分风格各有差异

大型语言模型(LLMs)越来越多地被用于生成科学评审意见,但现有评估很少在相同受控环境下检验不同提供商是否与会议决策及人类评审优先级保持一致。本研究对比了OpenAI GPT-5.4、Google Gemini 3.1 Pro Preview和Anthropic Claude Opus 4.6三种模型与人工评审意见及最终决策结果,基于300篇主题匹配的ICLR 2026投稿,均等分配至口头报告、海报和拒稿类别。每种模型在移除决策信息后,使用相同指令和评分标准评审所有论文。研究贡献在于跨提供商分析三个互补维度:与广泛和细粒度决策类别的对齐程度、推荐评分尺度的差异,以及识别弱点的主题一致性。所有三种模型均能区分接受与拒绝稿件,但均未能复现人类评审中口头报告与海报之间的区别。评分模式具有模型特异性:Gemini给出的评分系统性偏高,而OpenAI与Claude在拒稿和海报论文上更接近人类,但在口头报告上更为严苛。人类与模型评审在关注点上也存在差异:模型更频繁指出缺少基线对比,而人类更常提出计算效率问题。结果表明,广泛决策对齐并不意味着在精细人类判断或评审优先级上的共识。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.

大模型审稿学术评价人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。