arXiv:2509.21679cs.CL2025-09综述被引 3

用大模型检测学术评审中的错误论断,识别低质评论。

ReviewScore: Misinformed Peer Review Detection with Large Language Models

  • 从评审意见中提取隐含前提,自动判断其真假
  • 15.2%的缺点和26.4%的问题存在事实错误
  • 适合需要提升审稿质量的研究者与会议组织者

同行评审是学术研究的基石,但随着投稿量激增,人工智能会议中的评审质量持续下降。本文将‘误判性评审点’定义为包含错误前提的‘缺点’或论文已回答的‘问题’。我们验证了15.2%的缺点和26.4%的问题属于误判,并提出ReviewScore来标识评审点是否误判。为评估缺点中每个前提的事实性,我们设计了一个自动化引擎,重构显性和隐性前提。构建了人工专家标注的ReviewScore数据集,测试8个主流大模型在该任务上的表现。结果显示,模型的F1值为0.4–0.5,卡帕系数为0.3–0.4,表明中等一致但完全自动化仍具挑战。深入分析发现,多数错误源于模型推理偏差。此外,前提级事实性评估比整体缺点级评估显著提升人类与模型的一致性。

原文摘要 · Abstract (English)

Peer review serves as a backbone of academic research, but in most AI conferences, the review quality is degrading as the number of submissions explodes. To reliably detect low-quality reviews, we define misinformed review points as either "weaknesses" in a review that contain incorrect premises, or "questions" in a review that can be already answered by the paper. We verify that 15.2% of weaknesses and 26.4% of questions are misinformed and introduce ReviewScore indicating if a review point is misinformed. To evaluate the factuality of each premise of weaknesses, we propose an automated engine that reconstructs every explicit and implicit premise from a weakness. We build a human expert-annotated ReviewScore dataset to check the ability of LLMs to automate ReviewScore evaluation. Then, we measure human-model agreements on ReviewScore using eight current state-of-the-art LLMs. The models show F1 scores of 0.4--0.5 and kappa scores of 0.3--0.4, indicating moderate agreement but also suggesting that fully automating the evaluation remains challenging. A thorough disagreement analysis reveals that most errors are due to models' incorrect reasoning. We also prove that evaluating premise-level factuality shows significantly higher agreements than evaluating weakness-level factuality.

评审系统大模型事实核查AI会议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。