评测多模态深度研究智能体的端到端能力,强调图文证据一致性。
MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents
- 构建140个跨21领域的专家任务,含图文输入评估多模态理解与引用生成。
- 25个主流模型测试显示:文笔好不等于引用准,图文一致仍是瓶颈。
- 提出三阶段评估框架,可诊断报告质量、引用对齐与图文完整性问题。
深度研究智能体(DRAs)通过多步搜索与合成生成带引用的报告,但现有基准主要针对纯文本或短格式多模态问答,缺乏对端到端多模态证据使用的评估。本文提出MMDeepResearch-Bench(MMDR-Bench),包含140个专家设计的任务,覆盖21个领域,每项任务提供图像-文本组合,用于评估多模态理解与基于引用的报告生成能力。相比以往设置,MMDR-Bench强调以报告形式进行合成,明确要求模型将视觉元素与来源声明关联,并保持叙述、引用与视觉参考的一致性。我们进一步提出统一且可解释的评估流程:公式-大模型自适应评估(FLAE)用于报告质量,可信检索对齐引用评估(TRACE)用于引用与证据对齐,多模态支持对齐完整性检查(MOSAIC)用于文本-视觉一致性,各生成细粒度信号以支持错误诊断。在25个前沿模型上的实验揭示生成质量、引用规范性与多模态锚定之间的系统性权衡,表明优秀文笔并不保证忠实证据使用,多模态完整性仍是深度研究智能体的关键挑战。
原文摘要 · Abstract (English)
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where each task provides an image-text bundle to evaluate multimodal understanding and citation-grounded report generation. Compared to prior setups, MMDR-Bench emphasizes report-style synthesis with explicit evidence use, where models must connect visual artifacts to sourced claims and maintain consistency across narrative, citations, and visual references. We further propose a unified, interpretable evaluation pipeline: Formula-LLM Adaptive Evaluation (FLAE) for report quality, Trustworthy Retrieval-Aligned Citation Evaluation (TRACE) for citation-grounded evidence alignment, and Multimodal Support-Aligned Integrity Check (MOSAIC) for text-visual integrity, each producing fine-grained signals that support error diagnosis beyond a single overall score. Experiments across 25 state-of-the-art models reveal systematic trade-offs between generation quality, citation discipline, and multimodal grounding, highlighting that strong prose alone does not guarantee faithful evidence use and that multimodal integrity remains a key bottleneck for deep research agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。