arXiv:2506.02454cs.CLcs.AI2025-06AAAI被引 19

让AI从零生成图文混排的研究报告,自动设计图表并融合文本。

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

  • 用结构化文本描述图表,教会AI生成高质量可视化
  • 分四步构建框架,实现图文协同生成,整体胜率82%
  • 适合需要自动化报告生成的科研、商业分析场景

可视化在信息传达中至关重要。近年来,推理与检索增强生成技术使大语言模型具备深度研究和生成综合报告的能力。然而,现有框架多聚焦纯文本生成,对图文混合报告的自动生成研究不足。该任务面临设计有效图表并将其与文本无缝融合的挑战。为此,我们提出形式化图表描述(FDV),一种结构化的图表文本表示方法,使LLM能学习并生成多样且高质量的可视化内容。基于此表示,我们构建了多模态深度研究者(Multimodal DeepResearcher)代理框架,将任务分解为四个阶段:(1)研究,(2)示例报告文本化,(3)规划,(4)多模态报告生成。为评估生成结果,我们开发了多模态报告基准(MultimodalReportBench),包含100个不同主题作为输入,以及5项专用评估指标。在多个模型和评估方法上的实验表明,Multimodal DeepResearcher效果显著。值得注意的是,在使用相同Claude 3.7 Sonnet模型的情况下,其整体胜率相较基线方法提升至82%。

原文摘要 · Abstract (English)

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep research frameworks primarily focus on generating text-only content, leaving the automated generation of interleaved texts and visualizations underexplored. This novel task poses key challenges in designing informative visualizations and effectively integrating them with text reports. To address these challenges, we propose Formal Description of Visualization (FDV), a structured textual representation of charts that enables LLMs to learn from and generate diverse, high-quality visualizations. Building on this representation, we introduce Multimodal DeepResearcher, an agentic framework that decomposes the task into four stages: (1) researching, (2) exemplar report textualization, (3) planning, and (4) multimodal report generation. For the evaluation of generated multimodal reports, we develop MultimodalReportBench, which contains 100 diverse topics served as inputs along with 5 dedicated metrics. Extensive experiments across models and evaluation methods demonstrate the effectiveness of Multimodal DeepResearcher. Notably, utilizing the same Claude 3.7 Sonnet model, Multimodal DeepResearcher achieves an 82% overall win rate over the baseline method.

多模态生成智能报告图表生成LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。