arXiv:2607.14713econ.GNcs.CL2026-07

多智能体辩论并未提升AI对论文的反馈质量,作者更偏好单次生成结果。

Does Multi-Agent Debate Improve AI Feedback on Research Papers?

  • 用单次生成与两种自建多智能体辩论工具对比评估反馈效果
  • 作者偏好单次生成,优势达0.66分(95%置信区间0.32~1.00)
  • 警惕用AI评判AI,可能反向误导研究者判断

在一项预注册、匿名化的论文内实验中,44篇经济学元分析论文的作者对三种AI报告按改进论文的有用性进行排名:一种前沿模型的单次生成,以及我们构建的两种多智能体辩论工具。所有报告长度和模板一致。作者更倾向于单次生成,得分比mad-research高0.66分(95%置信区间0.32至1.00),比paper-workshop高0.57分(0.16至0.95),尽管paper-workshop消耗了约三十倍的计算资源。回忆期刊审稿意见的作者通常将其排第一,从不排最后;另一项独立评估中,三位AI裁判几乎总是将真实审稿意见排在最后。有趣的是,未参与任何报告生成的Gemini模型却会将paper-workshop排第一,逆转作者偏好。该反转警示我们不应以AI裁判替代作者判断。本研究衡量的是已完成论文的感知实用性,是否应由AI充当审稿人则是另一个问题。

原文摘要 · Abstract (English)

Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.

AI审稿多智能体实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。