arXiv:2601.07611cs.AI2026-01被引 5

用多智能体模拟专家评审,自动诊断论文真实且关键的弱点。

DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning

  • 构建角色化评审智能体,基于人类评审标准精准识别问题
  • 引入作者反驳机制,验证并优化弱点真实性,避免误判
  • 学习真实审稿习惯,按严重程度排序输出前K个核心缺陷

现有论文弱点识别方法存在三大不足:多智能体系统仅表面模拟人类角色,缺乏深层评审标准;忽视审稿偏见与作者反驳对质量验证的关键作用;输出无优先级的弱缺点列表。为此,本文提出DIAGPaper,一个三模块协同的多智能体框架。定制器模块依据人工定义的评审标准,生成具备特定专长的评审智能体;反驳模块引入作者智能体,与评审智能体进行结构化辩论,以验证和修正提出的弱点;优先级模块基于大规模人类审稿数据,学习评估已验证弱点的严重性,并向用户推送最严重的前K个问题。在AAAR与ReviewCritique两个基准上的实验表明,DIAGPaper显著优于现有方法,能生成更真实、更针对具体论文的弱点,并以用户友好方式排序呈现。

原文摘要 · Abstract (English)

Paper weakness identification using single-agent or multi-agent LLMs has attracted increasing attention, yet existing approaches exhibit key limitations. Many multi-agent systems simulate human roles at a surface level, missing the underlying criteria that lead experts to assess complementary intellectual aspects of a paper. Moreover, prior methods implicitly assume identified weaknesses are valid, ignoring reviewer bias, misunderstanding, and the critical role of author rebuttals in validating review quality. Finally, most systems output unranked weakness lists, rather than prioritizing the most consequential issues for users. In this work, we propose DIAGPaper, a novel multi-agent framework that addresses these challenges through three tightly integrated modules. The customizer module simulates human-defined review criteria and instantiates multiple reviewer agents with criterion-specific expertise. The rebuttal module introduces author agents that engage in structured debate with reviewer agents to validate and refine proposed weaknesses. The prioritizer module learns from large-scale human review practices to assess the severity of validated weaknesses and surfaces the top-K severest ones to users. Experiments on two benchmarks, AAAR and ReviewCritique, demonstrate that DIAGPaper substantially outperforms existing methods by producing more valid and more paper-specific weaknesses, while presenting them in a user-oriented, prioritized manner.

论文诊断多智能体评审分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。