arXiv:2510.15253cs.CLcs.CV2025-10ACL综述被引 8

解决文档理解中多模态信息融合难题,提升大模型推理能力

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding

  • 提出多模态RAG框架,统一处理文本、表格、图表等异构信息
  • 构建涵盖领域、检索粒度的分类体系,梳理关键数据集与应用案例
  • 适合研究文档智能、多模态生成的开发者与工程师参考

文档理解在金融分析到科学发现等应用中至关重要。现有方法或依赖OCR流水线输入大语言模型(LLM),或使用原生多模态大模型(MLLM),均存在局限:前者丢失结构细节,后者难以建模上下文。检索增强生成(RAG)可将模型锚定在外部数据,但文档的多模态特性——包含文本、表格、图表和版式——要求更先进的范式:多模态RAG。该方法实现跨所有模态的全局检索与推理,释放文档的完整智能潜力。本文系统综述了面向文档理解的多模态RAG。我们基于领域、检索模态与粒度提出分类体系,回顾图结构与代理框架等进展。同时汇总关键数据集、基准测试、应用场景及产业部署,并指出效率、细粒度表征与鲁棒性等开放挑战,为文档AI未来发展提供路线图。

原文摘要 · Abstract (English)

Document understanding is critical for applications from financial analysis to scientific discovery. Current approaches, whether OCR-based pipelines feeding Large Language Models (LLMs) or native Multimodal LLMs (MLLMs), face key limitations: the former loses structural detail, while the latter struggles with context modeling. Retrieval-Augmented Generation (RAG) helps ground models in external data, but documents' multimodal nature, i.e., combining text, tables, charts, and layout, demands a more advanced paradigm: Multimodal RAG. This approach enables holistic retrieval and reasoning across all modalities, unlocking comprehensive document intelligence. Recognizing its importance, this paper presents a systematic survey of Multimodal RAG for document understanding. We propose a taxonomy based on domain, retrieval modality, and granularity, and review advances involving graph structures and agentic frameworks. We also summarize key datasets, benchmarks, applications and industry deployment, and highlight open challenges in efficiency, fine-grained representation, and robustness, providing a roadmap for future progress in document AI.

文档理解多模态RAG大模型信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。