arXiv:2604.14261cs.CLcs.AI2026-04ACL综述被引 2

用评分标准和文献定位提升AI评审深度,让反馈更具体可信。

ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

论文配图:ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents
图 1 · 摘自论文原文
  • 分两阶段生成:先写初稿再用工具检索证据补充
  • 在8个维度上优于大模型基线,连GPT-4.1也落后
  • 适合需要高质量评审的学术会议或期刊编辑

AI论文投稿激增促使研究者探索大语言模型辅助同行评审。但现有LLM评审常生成表面化、模板化评论,缺乏实质性和证据支撑。我们归因于未充分利用人类评审的两大关键要素:明确评分标准和对已有工作的上下文定位。为此,我们构建REVIEWBENCH基准,基于论文内容、官方指南与人工评审提取专属评分标准,评估评审文本质量。进一步提出REVIEWGROUNDER框架,采用评分引导、工具集成的多智能体机制,将评审分解为起草与证据定位两个阶段,通过定向证据整合增强浅层草稿。在REVIEWBENCH上的实验表明,使用Phi-4-14B作起草、GPT-OSS-120B作证据定位的REVIEWGROUNDER,在8个维度上均显著优于基线,包括性能更强的大模型如GPT-4.1与DeepSeek-R1-670B,且更贴近人类评审判断。代码已开源。

原文摘要 · Abstract (English)

The rapid rise in AI conference submissions has driven increasing exploration of large language models (LLMs) for peer review support. However, LLM-based reviewers often generate superficial, formulaic comments lacking substantive, evidence-grounded feedback. We attribute this to the underutilization of two key components of human reviewing: explicit rubrics and contextual grounding in existing work. To address this, we introduce REVIEWBENCH, a benchmark evaluating review text according to paper-specific rubrics derived from official guidelines, the paper's content, and human-written reviews. We further propose REVIEWGROUNDER, a rubric-guided, tool-integrated multi-agent framework that decomposes reviewing into drafting and grounding stages, enriching shallow drafts via targeted evidence consolidation. Experiments on REVIEWBENCH show that REVIEWGROUNDER, using a Phi-4-14B-based drafter and a GPT-OSS-120B-based grounding stage, consistently outperforms baselines with substantially stronger/larger backbones (e.g., GPT-4.1 and DeepSeek-R1-670B) in both alignment with human judgments and rubric-based review quality across 8 dimensions. The code is available \href{https://github.com/EigenTom/ReviewGrounder}{here}.

AI评审多智能体评分标准证据定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。