arXiv:2605.13034cs.CVcs.IR2026-05

让论文图表真正成为研究报告的证据,提升可验证性。

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

论文配图:ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence
图 1 · 摘自论文原文
  • 将源图当作可检索、可解析的证据单元,精准关联论点。
  • 通过上下文过滤和视觉分析,把杂乱网页图转化为可信证据。
  • 适合需要高可信度报告的科研与政策分析场景。

近期深度研究系统通过迭代检索与推理提升了大语言模型生成长篇有依据报告的能力。然而,多数文本主导系统仅依赖文字证据,多模态系统常弱化图像检索或自行生成图表,导致源图证据未被充分利用。我们提出ViDR,一种将长篇报告扎根于源图证据的多模态深度研究框架。ViDR将源图视为可检索、可解释、可路由、可验证的证据对象,同时在需要时生成分析图表。它构建基于证据的提纲,将论点与文本和视觉证据关联;通过上下文感知过滤、提纲感知重排和视觉语言模型(VLM)分析,将噪声网页图净化为源图证据原子;并为每个章节生成特定证据支持的内容。ViDR还验证视觉引用以减少幻觉或错位图像。我们还引入MMR Bench+,用于评估研究报告中视觉证据的使用情况,涵盖源图检索、定位、解读、可验证性及分析图表生成。实验表明,相比强商业与开源基线,ViDR在整体报告质量、源图整合度和可验证性上均有提升。结果表明,源视觉证据对多模态深度研究至关重要,能增强证据基础、视觉支持与报告可验证性。

原文摘要 · Abstract (English)

Recent deep research systems have improved the ability of large language models to produce long, grounded reports through iterative retrieval and reasoning. However, most text-centered systems rely mainly on textual evidence, while multimodal systems often retrieve images only weakly or generate charts themselves, leaving source figures underused as evidence. We present ViDR, a multimodal deep research framework that grounds long-form reports in source figures. ViDR treats source figures as retrievable, interpretable, routable, and verifiable evidence objects, while still generating analytical charts when needed. It builds an evidence-indexed outline linking claims to textual and visual evidence, refines noisy web images into source-figure evidence atoms through context-aware filtering, outline-aware reranking, and VLM-based visual analysis, and generates each section with section-specific evidence. ViDR further validates visual references to reduce hallucinated or misplaced figures. We also introduce MMR Bench+, a benchmark for evaluating visual evidence use in deep research reports, covering source-figure retrieval, placement, interpretation, verifiability, and analytical chart generation. Experiments show that ViDR improves overall report quality, source-figure integration, and verifiability over strong commercial and open-source baselines. These results suggest that source visual evidence is important for multimodal deep research, as it strengthens evidential grounding, visual support, and report verifiability.

多模态报告生成视觉证据可验证性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。