arXiv:2601.03262cs.IRcs.CL2026-01中稿 · AACL-IJCNLP 2025综述被引 2

MLLM如何提升图文混排文档在RAG中的检索效果

Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey

  • 将MLLM分为统一描述、多模态嵌入和端到端表示三类角色
  • 不同角色在精度、速度、索引大小上各有优劣
  • 适合需要高精度图文联合检索的系统开发者参考

图文混排文档(VRDs)因版面语义依赖、光学字符识别脆弱以及信息分散在复杂图表与表格中,给检索增强生成(RAG)带来挑战。本文综述多模态大模型(MLLMs)在解决此类问题中的应用,将现有研究归纳为三类角色:模态统一描述器、多模态嵌入器与端到端表示器。从检索粒度、信息保真度、延迟与索引大小等方面比较各角色表现,并讨论其与重排序和定位机制的兼容性。文章还分析关键权衡,提供实际选型建议,并指出未来方向,包括自适应检索单元设计、模型压缩及评估方法构建。

原文摘要 · Abstract (English)

Visually rich documents (VRDs) challenge retrieval-augmented generation (RAG) with layout-dependent semantics, brittle OCR, and evidence spread across complex figures and structured tables. This survey examines how Multimodal Large Language Models (MLLMs) are being used to make VRD retrieval practical for RAG. We organize the literature into three roles: Modality-Unifying Captioners, Multimodal Embedders, and End-to-End Representers. We compare these roles along retrieval granularity, information fidelity, latency and index size, and compatibility with reranking and grounding. We also outline key trade-offs and offer some practical guidance on when to favor each role. Finally, we identify promising directions for future research, including adaptive retrieval units, model size reduction, and the development of evaluation methods.

多模态RAG文档检索MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。