评估大模型论文评审时,发现其忽视创新性而过度关注技术正确性。
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
- 基于人类专家标注的3657个优缺点,构建注意力分布评估框架。
- 676篇论文评审显示,大模型对新颖性的关注远低于人类专家。
- 适合研究AI评审可信度、提升自动化审稿质量的学者使用。
同行评审是科学进步的基础,但正面临审稿人短缺和工作量激增的压力。大语言模型(LLMs)可自动生成评审意见,但其可信度需系统评估。现有方法仅在表面层面(如BLEU、ROUGE)或内容层面(如具体性、事实准确性)进行评估,却未检验大模型是否关注人类专家决策中的关键维度——即决定论文接受与否的优缺点。本文提出一种聚焦级评估框架,将评审重点建模为预定义评审维度上的归一化注意力分布。基于该框架,我们构建了自动评估流水线,涵盖两类维度:目标(如问题、方法、实验)与方面(如有效性、清晰性、新颖性)。利用来自OpenReview的676篇论文评审(共3,657个优缺点),对比大模型与人类专家的注意力分布,发现现成大模型在批评论文时持续偏向技术有效性,显著忽视新颖性评估。
原文摘要 · Abstract (English)
Peer review underpins scientific progress, but it is increasingly strained by reviewer shortages and growing workloads. Large Language Models (LLMs) can automatically draft reviews now, but determining whether LLM-generated reviews are trustworthy requires systematic evaluation. Researchers have evaluated LLM reviews at either surface-level (e.g., BLEU and ROUGE) or content-level (e.g., specificity and factual accuracy). Yet it remains uncertain whether LLM-generated reviews attend to the same critical facets that human experts weigh -- the strengths and weaknesses that ultimately drive an accept-or-reject decision. We introduce a focus-level evaluation framework that operationalizes the focus as a normalized distribution of attention across predefined facets in paper reviews. Based on the framework, we developed an automatic focus-level evaluation pipeline based on two sets of facets: target (e.g., problem, method, and experiment) and aspect (e.g., validity, clarity, and novelty), leveraging 676 paper reviews (https://figshare.com/s/d5adf26c802527dd0f62) from OpenReview that consists of 3,657 strengths and weaknesses identified from human experts. The comparison of focus distributions between LLMs and human experts showed that the off-the-shelf LLMs consistently have a more biased focus towards examining technical validity while significantly overlooking novelty assessment when criticizing papers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。