arXiv:2609.09076cs.CL2026-09综述

让AI根据作者回复生成可操作的修改建议,提升审稿反馈质量。

ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation

论文配图:ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
图 1 · 摘自论文原文
  • 基于作者回复挖掘修改动作,构建诊断与建议的对应关系。
  • 在1000个实例上验证,修改建议实用性显著优于现有模型。
  • 适合希望提升论文修改效率的研究者和审稿人使用。

随着大语言模型被广泛用于投稿前自审,人们越来越需要不仅能指出问题、还能引导具体修改的反馈。本文将此任务拆解为诊断性观点生成与修改建议生成两个子任务,提出ActReview:一种基于作者回复的后训练框架,将论文特定诊断与具体、有依据的修改方案相连接。核心思路是:作者对审稿意见的回应揭示了可行的修改路径,可作为修改导向反馈的潜在监督信号。我们从OpenReview的真实审稿-回复对话中构建了包含40,000条数据的ActReview-40K数据集,将审稿弱点与作者回应对齐,并将反馈锚定在论文局部证据上。我们使用多任务监督微调和候选感知、弱点特异性评分器的GRPO方法,对Qwen3-8B-Base进行后训练。同时构建了人工标注的1,000实例基准测试集ActReview-Bench,用于评估诊断质量和修改实用性。实验表明,ActReview在可操作性和依据性上优于先前专用模型,且与强提示工程大模型相当。人工评估证实修改建议实用性提升,但技术准确性仍有不足;额外分析支持其在未见论文上的泛化能力及跨评审者的鲁棒性。

原文摘要 · Abstract (English)

As LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identifies weaknesses but also guides authors toward concrete revisions. We study this as Actionable Peer-review Generation and decompose it into two subtasks: diagnostic claim generation and revision suggestion generation. We introduce ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans. Our central insight is that author rebuttals reveal plausible actions for addressing reviewer concerns and can therefore provide latent supervision for revision-oriented feedback. From real review-rebuttal threads on OpenReview, we construct ActReview-40K by aligning reviewer weaknesses with author responses and grounding the resulting feedback in localized paper evidence. We post-train Qwen3-8B-Base with multi-task supervised fine-tuning followed by GRPO using candidate-aware, weakness-specific rubric rewards. We also introduce ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness. Experiments show that ActReview outperforms prior specialized review-generation models on actionability and grounding while remaining competitive with strong prompt-based LLMs. Human evaluation confirms improved revision usefulness while revealing a remaining gap in technical accuracy, and additional analyses support generalization to held-out papers and robustness across independent judges.

可操作反馈审稿生成大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。