让AI生成有图有据的多模态研究报告,自动验证事实与图文一致性。
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

- 用多个专业代理协同规划、搜证、写报告,图文同步
- 在视觉工作记忆中保存带来源标注的图像,保证可追溯
- 内置验证器确保内容真实、引用准确、图文一致,适合科研助手
大型语言模型已将自主智能体从获取简明事实答案的深度搜索,推进到整合零散证据生成长篇报告的深度研究。然而,可验证的多模态深度研究仍具挑战性,原因在于开放式合成缺乏确定性真值,且需将文本论点与视觉证据交织呈现。本文提出Ptah,一种用于交错式报告生成的多智能体协作框架。Ptah通过规划、研究与写作三阶段,协调用户查询至网页报告的全生命周期:专用代理构建视觉感知计划,收集基于主张的证据,在视觉工作记忆中保持源对齐的图像,并通过声明式多模态工具使用撰写报告。验证代理作为框架的接受函数,全程强制执行事实依据、引用一致性和跨模态一致性。我们进一步提出PtahEval评估协议,扩充现有基准以包含图像级和展示级评估。在深度研究基准上的实验表明,Ptah生成的报告比强基线更可靠、视觉信息更丰富、对人类更可用。代码已开源:https://github.com/SnowNation101/Ptah
原文摘要 · Abstract (English)
Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports. However, verifiable multimodal deep research remains challenging due to open-ended synthesis without deterministic ground truth and the need to interleave textual arguments with visual evidence. We propose Ptah, a multi-agent harness for interleaved report generation. Ptah orchestrates the lifecycle from user query to rendered web report through planning, research, and writing stages, where specialized agents construct visual-aware plans, collect claim-grounded evidence, maintain source-aligned images in a Visual Working Memory, and compose reports through declarative multimodal tool use. A verifier agent serves as the harness's acceptance function, enforcing factual grounding, citation fidelity, and cross-modal consistency throughout the workflow. We further introduce PtahEval, an evaluation protocol that augments existing benchmarks with image-level and presentation-level assessments. Experiments on deep research benchmarks show that Ptah produces more reliable, visually informative, and usable human-facing multimodal reports than strong baselines. Our code is released at https://github.com/SnowNation101/Ptah
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。