arXiv:2606.20897cs.CLcs.AI2026-06ACL被引 1

用PeerCheck框架提升大模型学术评审质量,让机器评阅更像人。

PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

论文配图:PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality
图 1 · 摘自论文原文
  • 通过对比人类与大模型评审差异,发现模型偏理论、人重方法实验
  • 链式思考提示显著提升评审质量,但检索增强生成反而可能降分
  • 适合想改进自动评审系统的研究者,尤其关注人机对齐的场景

随着学术投稿量增加,传统同行评审难以应对,质量与公平性受关注。使用大语言模型(LLMs)辅助成为趋势。本文提出PeerCheck框架,分析大模型与人类评审的差异(RQ1),探索提升大模型评审质量的方法(RQ2)。我们对比了不同大模型生成的评审与人工评审,发现大模型更关注理论,而人类更重视方法与实验。通过链式思考(CoT)提示和检索增强生成(RAG)改进评审质量,结果表明CoT显著提升质量,但出现意外的“RAG悖论”:不同大模型下效果不一,某些情况下甚至降低评审质量。本研究揭示了大模型评审的潜力与局限,推动更贴近人类标准的评审系统发展。数据集已开源于https://github.com/TrustAIRLab/PeerCheck。

原文摘要 · Abstract (English)

As academic submissions grow, the traditional peer review process struggles to keep up, raising concerns about quality and fairness. A trend of using large language models (LLMs) for assistance has emerged. In this work, we take a critical step toward improving the quality of LLM-generated reviews. We propose the PeerCheck framework, which investigates LLM-human review differences (RQ1) and explores methods to improve LLM-generated review quality (RQ2). We first analyzed the human-written reviews with reviews generated by various LLMs and found that LLMs and humans focus on different terms, e.g., LLMs prioritize theory while humans emphasize methodology and experiments. We further adopt prompt engineering, such as Chain-of-Thought (CoT), and utilize retrieval-augmented generation (RAG) to enhance the LLM-generated reviews towards human-level quality. We find CoT significantly improves the quality of LLM reviews, while we discover an unexpected "RAG paradox," i.e., experiments with RAG produce different results for various LLMs and, in some cases, even reduce review quality. Our comprehensive analysis of LLM-generated academic reviews illustrates both possibilities and limitations, contributing to a more effective, human-aligned review system. Our dataset is available on https://github.com/TrustAIRLab/PeerCheck.

大模型评审人机对齐链式思考RAG悖论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。