arXiv:2607.22553cs.CLcs.AI2026-07ACL综述被引 1

对比不同评审指南对AI评阅效果的影响,发现真实会议标准更优。

Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

论文配图:Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
图 1 · 摘自论文原文
  • 用大模型生成模仿人类高质量评审的指南
  • 官方会议指南比仿真人指南更贴近人工评判
  • 严格评分标准反而降低评阅质量,需保留主观判断

同行评审是科研中的关键环节,但日益增长的工作量使其自动化变得迫切。本研究分析了不同类型评审指南(如正式会议指南与由大模型基于高质量人工评审生成的仿真人指南)对自动化评审的影响。实验表明,官方会议指南产生的评审结果最接近人工判断,说明经过会议实践优化的评价标准同样适用于自动化评审。相比之下,仿真人指南总体效果较差。此外,强制执行严格的评分模板会持续降低性能,凸显了允许主观与整体性评分的重要性。

原文摘要 · Abstract (English)

Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.

自动化评审大模型评估标准会议指南

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。