arXiv:2510.21603cs.IRcs.CL2025-10

让大模型能像人一样深度研读图文并茂的文档

Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research

  • 用多粒度解析保留图表公式的视觉语义,构建文档级理解
  • 在158个复杂问题上达50.6%准确率,比现有方法快3.4倍
  • 适合需要跨文档、跨模态推理的研究者与知识工作者

深度研究系统通过迭代推理和证据收集革新了大模型解决复杂问题的方式。然而,现有系统仍局限于文本网页数据,忽视了多模态文档中蕴含的海量知识。处理此类文档需精确解析以保留视觉语义(如图表、表格、公式),智能分块维持结构连贯性,并支持跨模态自适应检索,而这些能力在现有系统中缺失。为此,我们提出Doc-Researcher,一个统一系统,包含三个核心组件:(i) 深度多模态解析,在保持布局结构的同时生成从片段到文档级别的多粒度表征;(ii) 系统化检索架构,支持纯文本、纯视觉及混合范式,并动态选择粒度;(iii) 迭代多智能体工作流,分解复杂查询,逐步积累证据,并跨文档与模态合成完整答案。为实现严格评估,我们引入M4DocBench——首个面向多模态、多跳、多文档、多轮次深度研究的基准。该基准包含158个专家标注的问题,覆盖304份文档,完整证据链,测试现有基准无法评估的能力。实验表明,Doc-Researcher达到50.6%准确率,比最先进基线高出3.4倍,验证了有效文档研究不仅依赖更好检索,更需深层解析以保持多模态完整性并支持迭代研究。本工作确立了对多模态文档集合进行深度研究的新范式。

原文摘要 · Abstract (English)

Deep Research systems have revolutionized how LLMs solve complex questions through iterative reasoning and evidence gathering. However, current systems remain fundamentally constrained to textual web data, overlooking the vast knowledge embedded in multimodal documents Processing such documents demands sophisticated parsing to preserve visual semantics (figures, tables, charts, and equations), intelligent chunking to maintain structural coherence, and adaptive retrieval across modalities, which are capabilities absent in existing systems. In response, we present Doc-Researcher, a unified system that bridges this gap through three integrated components: (i) deep multimodal parsing that preserves layout structure and visual semantics while creating multi-granular representations from chunk to document level, (ii) systematic retrieval architecture supporting text-only, vision-only, and hybrid paradigms with dynamic granularity selection, and (iii) iterative multi-agent workflows that decompose complex queries, progressively accumulate evidence, and synthesize comprehensive answers across documents and modalities. To enable rigorous evaluation, we introduce M4DocBench, the first benchmark for Multi-modal, Multi-hop, Multi-document, and Multi-turn deep research. Featuring 158 expert-annotated questions with complete evidence chains across 304 documents, M4DocBench tests capabilities that existing benchmarks cannot assess. Experiments demonstrate that Doc-Researcher achieves 50.6% accuracy, 3.4xbetter than state-of-the-art baselines, validating that effective document research requires not just better retrieval, but fundamentally deep parsing that preserve multimodal integrity and support iterative research. Our work establishes a new paradigm for conducting deep research on multimodal document collections.

文档理解多模态深度研究多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。