arXiv:2606.02320cs.CL2026-06被引 1

构建可生成图文混排报告的智能研究代理,提升分析可信度。

TVIR: Building Deep Research Agents Towards Text-Visual Interleaved Report Generation

论文配图:TVIR: Building Deep Research Agents Towards Text-Visual Interleaved Report Generation
图 1 · 摘自论文原文
  • 设计分层多智能体框架,协同完成大纲、图像检索与图表生成。
  • 在100项专家标注任务中表现优异,图文对齐度显著提升。
  • 适合需要高可信度图文报告的研究者与AI系统开发者。

深度研究代理在多步信息检索、推理和长文本报告生成方面表现出强大能力,但现有基准与系统仍以文本为中心,缺乏对视觉元素事实可靠性及其与分析内容一致性的评估。为弥补这一空白,我们提出TVIR(文本-视觉交错报告生成),包含TVIR-Bench——一个由100个专家精心设计的多模态深度研究任务集合,要求视觉元素服务于特定分析子目标;以及TVIR-Agent——一种分层多智能体框架,可作为构建提纲、检索图像、生成可追溯来源的图表,并通过上下文感知的顺序写作整合报告的强基线模型。我们进一步开发了双路径评估框架,结合文本评估与视觉评估。在九个深度研究系统上的实验表明,TVIR-Agent整体性能出色,凸显了显式多模态设计与评估在证据驱动报告生成中的重要性。

原文摘要 · Abstract (English)

Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmarks and systems remain predominantly text-centric, with limited evaluation of whether visual elements are factually reliable and well aligned with the surrounding analysis. To address this gap, we introduce TVIR (Text-Visual Interleaved Report Generation), which includes TVIR-Bench, a benchmark of 100 expert-curated multimodal deep research tasks that require visual elements to serve specific analytical sub-goals, and TVIR-Agent, a hierarchical multi-agent framework that serves as a strong baseline for constructing outlines, retrieving images, generating charts with traceable sources, and composing reports through context-aware sequential writing. We further develop a dual-path evaluation framework that combines Textual Assessment and Visual Assessment. Experiments across nine deep research systems show that TVIR-Agent achieves strong overall performance, underscoring the importance of explicit multimodal design and evaluation for evidence-driven report generation.

多模态智能代理报告生成图文对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。