综述多模态大模型在图文文档理解中的方法与挑战
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
- 梳理文本、图像、版式特征的融合表示技术
- 分析预训练、指令微调等核心训练范式
- 适合关注文档智能与多模态模型的研究者
图文丰富文档理解(VRDU)因需自动解析包含复杂视觉、文本和结构元素的文档而成为研究重点。近年来,多模态大语言模型(MLLMs)在该领域展现出显著潜力,涵盖基于OCR与无OCR的信息抽取方法。本文综述了基于MLLM的VRDU最新进展,聚焦两大关键方面:(1) 文本、视觉与版式特征的表征与融合技术;(2) 预训练、指令微调及训练策略。同时,探讨数据稀缺、多页与多语言文档处理等挑战,并融合检索增强生成与代理框架等新兴趋势。分析为构建更可扩展、可靠与自适应的MLLM驱动的VRDU系统提供路线图。
原文摘要 · Abstract (English)
Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, and structural elements. Recently, Multimodal Large Language Models (MLLMs) have demonstrated significant promise in this domain, including both OCR-based and OCR-free approaches for information extraction from document images. This survey reviews recent advances in MLLM-based VRDU, highlighting emerging trends and promising research directions with a focus on two key aspects: (1) techniques for representing and integrating textual, visual, and layout features; (2) training paradigms, including pretraining, instruction tuning, and training strategies. Moreover, we address challenges such as data scarcity, handling multi-page and multilingual documents, and integrating emerging trends such as Retrieval-Augmented Generation and agentic frameworks. Our analysis offers a roadmap for advancing MLLM-based VRDU toward more scalable, reliable, and adaptable systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。