arXiv:2608.00005cs.CLcs.AI2026-08综述

用显式评分标准提升论文评审的全面性与可靠性

RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review

论文配图:RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review
图 1 · 摘自论文原文
  • 先生成适配论文的评分标准,再据此生成评审意见
  • 比现有系统更全面、更具区分度,抗攻击能力最强
  • 适合需要高质量评审的学术会议和期刊

大型会议面临空前的投稿压力,促使使用大语言模型(LLMs)作为评审助手。现有基于LLM的评审系统存在两个结构性缺陷:一是将论文直接映射为评审意见,使评分标准隐含且难以分离;二是现有范式仅体现评审的一半特征——无训练代理虽能收集广泛证据但缺乏方向性,而有监督模型虽继承人类判别力却也携带噪声与覆盖不均。我们提出RubricReviewer,一个完全基于评分标准的框架,将评分标准生成作为显式中间步骤,使评审生成与最终评估均依赖于论文自适应的评分标准。该框架结合无训练代理(Scout)获取外部证据与经人类对齐的训练模型(Aligner)处理这些证据,融合两类监督来源的优势。在真实投稿数据上的实验表明,RubricReviewer生成的评审意见显著更全面、更具区分度,且对对抗性提示注入攻击表现出最强鲁棒性。消融实验进一步验证了各组件的必要性。

原文摘要 · Abstract (English)

Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants. Existing LLM-based reviewers, however, face two structural limitations. First, they map manuscripts directly to reviews, leaving the underlying rubric implicit and entangling its derivation with the judgement. Second, the prevailing paradigms each capture only half of a good review: training-free agents gather broad evidence but produce undirected critiques, while training-based reviewers inherit human discriminative judgement together with its noise and uneven coverage. We introduce RubricReviewer, a fully rubric-driven framework that addresses both limitations. It makes rubric generation an explicit intermediate step, so that both review generation and the final assessment are conditioned on paper-adaptive rubrics. It further combines a training-free agent (Scout) that gathers external evidence with a human-aligned trained model (Aligner) that consumes this evidence, fusing the strengths of both supervision sources. Experiments on real-world submissions show that RubricReviewer produces reviews that are markedly more comprehensive and more discriminative than prior systems, and exhibits the strongest robustness against adversarial prompt-injection attacks. Ablation studies further confirm the necessity of each component.

论文评审大模型评分标准自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。