现有论文生成工具在反复修改时会丢内容,可靠性存疑。
Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision

- 用多轮反馈模拟人类修订流程,测试论文生成工具的迭代能力
- 五款工具修改时16%-27%已写内容出错,越改越乱
- 单纯优化提示词或加子任务模块也解决不了根本问题
当前深度研究代理(DRAs)的评测将报告生成视为单次写作任务,这与人类研究者通过自我反思或同行反馈反复修改报告的实际过程严重不符。现有方法未考察代理在用户反馈下是否能可靠地进行多轮修订。为此,我们提出 Mr Dre 评估框架,建立多轮报告修订新评测维度:(1)统一的长文本报告评估协议,涵盖全面性、事实性和呈现质量;(2)基于人工验证的反馈模拟流水线,支持多轮修订。对五种不同 DRAs 的分析表明,尽管代理能响应大多数用户反馈,却会在 16%-27% 的已覆盖内容和引文质量上出现退化。在多轮修订中,即使表现最好的代理仍有显著提升空间,持续破坏反馈范围外的内容,并无法保留早期修改成果。进一步实验显示,这些缺陷难以通过推理阶段的提示工程或专用修订子代理等方法缓解。
原文摘要 · Abstract (English)
Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task, which fundamentally diverges from how human researchers iteratively draft and revise reports via self-reflection or peer feedback. Whether DRAs can reliably revise reports with user feedback remains unexplored. We introduce Mr Dre, an evaluation suite that establishes multi-turn report revision as a new evaluation axis for DRAs. Mr Dre consists of (1) a unified long-form report evaluation protocol spanning comprehensiveness, factuality, and presentation, and (2) a human-verified feedback simulation pipeline for multi-turn revision. Our analysis of five diverse DRAs reveals a critical limitation: while agents can address most user feedback, they also regress on 16-27% of previously covered content and citation quality. Over multiple revision turns, even the best-performing agents leave significant headroom, as they continue to disrupt content outside the feedback's scope and fail to preserve earlier edits. We further show that these issues are not easily resolvable through inference-time fixes such as prompt engineering and a dedicated sub-agent for report revision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。