arXiv:2603.27360cs.AI2026-03综述被引 3

用AI自动生成审稿回复,作者只需少量引导。

Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance

  • 让作者仅通过简单指引驱动推理,实现高效反驳生成。
  • 相比直接用大模型生成,事实准确率提升显著。
  • 适合希望快速回应审稿意见的科研人员使用。

反驳生成是科学论文同行评审过程中的关键环节,使作者能够澄清误解、纠正事实错误,并引导审稿人做出更准确的评估。我们发现,直接使用大语言模型(LLMs)进行反驳生成时,往往难以实现针对性反驳且保持事实准确性,凸显出结构化推理和作者干预的必要性。为此,本文提出DEFEND——一种基于大语言模型的工具,旨在显式执行自动化反驳生成背后的推理过程,同时保持作者在环路中。与从零开始撰写反驳不同,作者只需以最少干预驱动推理流程,从而实现低负担、高效率的反驳生成。我们对比了四种范式:(i) 直接使用LLM生成反驳(DRG),(ii) 分段式使用LLM生成反驳(SWRG),(iii) 无作者干预的分段式自动序列方法(SA)。为实现细粒度评估,我们扩展了ReviewCritique数据集,新增审稿意见分段、缺陷标注、错误类型、反驳动作标签及与黄金反驳段的映射。实验结果与用户研究显示,直接使用LLM在事实正确性和针对性反驳上表现较差;而分段生成与带作者参与的自动序列方法显著提升了事实准确性和反驳力度。

原文摘要 · Abstract (English)

Rebuttal generation is a critical component of the peer review process for scientific papers, enabling authors to clarify misunderstandings, correct factual inaccuracies, and guide reviewers toward a more accurate evaluation. We observe that Large Language Models (LLMs) often struggle to perform targeted refutation and maintain accurate factual grounding when used directly for rebuttal generation, highlighting the need for structured reasoning and author intervention. To address this, in the paper, we introduce DEFEND an LLM based tool designed to explicitly execute the underlying reasoning process of automated rebuttal generation, while keeping the author-in-the-loop. As opposed to writing the rebuttals from scratch, the author needs to only drive the reasoning process with minimal intervention, leading an efficient approach with minimal effort and less cognitive load. We compare DEFEND against three other paradigms: (i) Direct rebuttal generation using LLM (DRG), (ii) Segment-wise rebuttal generation using LLM (SWRG), and (iii) Sequential approach (SA) of segment-wise rebuttal generation without author intervention. To enable finegrained evaluation, we extend the ReviewCritique dataset, creating review segmentation, deficiency, error type annotations, rebuttal-action labels, and mapping to gold rebuttal segments. Experimental results and a user study demonstrate that directly using LLMs perform poorly in factual correctness and targeted refutation. Segment-wise generation and the automated sequential approach with author-in-the-loop, substantially improve factual correctness and strength of refutation.

AI辅助审稿回复大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。