让AI写带图文的长报告,更准更连贯。
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation

- 用智能体迭代搜文本和图像,边查边筛。
- 8000条高质量生成轨迹训练模型,提升准确性。
- 适合需要图文结合的科研、报告写作场景。
近期的智能体搜索框架通过迭代规划与检索实现深度研究,减少幻觉并增强事实依据。然而,这些方法仍以文本为中心,忽视了真实专家报告中常见的多模态证据。我们提出一项紧迫任务:多模态长篇生成。为此,我们构建 Deep-Reporter,一个统一的智能体框架,用于生成基于多模态证据的长篇内容。该框架包含:(i) 智能体多模态搜索与过滤,用于检索并筛选文本段落和信息密集型图像;(ii) 检查清单引导的渐进式合成,确保图文融合连贯且引用位置最优;(iii) 循环上下文管理,平衡长程连贯性与局部流畅性。我们建立了一个严谨的数据清洗流程,生成8000条高质量智能体轨迹用于模型优化。此外,我们还引入 M2LongBench,一个涵盖9个领域共247项研究任务的综合性测试平台,配备稳定的多模态沙盒环境。大量实验表明,多模态长篇生成是一项具有挑战性的任务,尤其在多模态选择与整合方面,有效的后训练可显著缩小性能差距。
原文摘要 · Abstract (English)
Recent agentic search frameworks enable deep research via iterative planning and retrieval, reducing hallucinations and enhancing factual grounding. However, they remain text-centric, overlooking the multimodal evidence that characterizes real-world expert reports. We introduce a pressing task: multimodal long-form generation. Accordingly, we propose Deep-Reporter, a unified agentic framework for grounded multimodal long-form generation. It orchestrates: (i) Agentic Multimodal Search and Filtering to retrieve and filter textual passages and information-dense visuals; (ii) Checklist-Guided Incremental Synthesis to ensure coherent image-text integration and optimal citation placement; and (iii) Recurrent Context Management to balance long-range coherence with local fluency. We develop a rigorous curation pipeline producing 8K high-quality agentic traces for model optimization. We further introduce M2LongBench, a comprehensive testbed comprising 247 research tasks across 9 domains and a stable multimodal sandbox. Extensive experiments demonstrate that long-form multimodal generation is a challenging task, especially in multimodal selection and integration, and effective post-training can bridge the gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。