arXiv:2601.15715cs.CLcs.AI2026-01中稿 · ICLR被引 3

用心理理论让AI更懂审稿人,写出更有说服力的论文回应。

RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind

  • 基于心理理论构建三步框架,模拟审稿人想法并制定回应策略。
  • 在自动与人工评估中均优于基线模型18.3%以上,超越部分商业模型。
  • 专为学术反驳设计,适合科研人员提升投稿回复质量。

尽管人工智能已深度融入研究流程并取得显著进展,学术反驳仍是一个复杂且未被充分探索的挑战。这是因为反驳本质上是在严重信息不对称下的战略沟通,而非简单的技术争论。当前方法多停留在表面语言模仿,缺乏有效说服所必需的换位思考能力。本文提出RebuttalAgent,首个将学术反驳建立在心理理论(ToM)基础上的框架,通过ToM-策略-回应(TSR)机制建模审稿人心理状态、制定说服策略并生成有证据支持的回应。为训练该代理,我们构建了大规模数据集RebuttalBench,采用创新的批判与优化方法合成。训练分为两阶段:先进行监督微调以赋予模型基于心理理论的分析与规划能力,再通过自奖励强化学习实现可扩展的自我提升。为实现可靠高效的自动化评估,我们进一步开发了Rebuttal-RM,一个在超10万条多源反驳数据上训练的专用评价器,其评分一致性超越强大裁判GPT-4.1。大量实验表明,RebuttalAgent在自动指标上平均领先基线模型18.3%,并在自动与人工评估中均优于先进商用模型。

原文摘要 · Abstract (English)

Although artificial intelligence (AI) has become deeply integrated into various stages of the research workflow and achieved remarkable advancements, academic rebuttal remains a significant and underexplored challenge. This is because rebuttal is a complex process of strategic communication under severe information asymmetry rather than a simple technical debate. Consequently, current approaches struggle as they largely imitate surface-level linguistics, missing the essential element of perspective-taking required for effective persuasion. In this paper, we introduce RebuttalAgent, the first framework to ground academic rebuttal in Theory of Mind (ToM), operationalized through a ToM-Strategy-Response (TSR) framework that models reviewer mental state, formulates persuasion strategy, and generates evidence-based response. To train our agent, we construct RebuttalBench, a large-scale dataset synthesized via a novel critique-and-refine approach. Our training process consists of two stages, beginning with a supervised fine-tuning phase to equip the agent with ToM-based analysis and strategic planning capabilities, followed by a reinforcement learning phase leveraging the self-reward mechanism for scalable self-improvement. For reliable and efficient automated evaluation, we further develop Rebuttal-RM, a specialized evaluator trained on over 100K samples of multi-source rebuttal data, which achieves scoring consistency with human preferences surpassing powerful judge GPT-4.1. Extensive experiments show RebuttalAgent significantly outperforms the base model by an average of 18.3% on automated metrics, while also outperforming advanced proprietary models across both automated and human evaluations.

AI写作心理理论学术投稿智能反驳

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。