arXiv:2504.08748cs.IRcs.AI2025-04综述被引 56

让大模型同时看图识文生成答案,减少幻觉。

A Survey of Multimodal Retrieval-Augmented Generation

  • 用图文视频联合检索增强语言模型生成能力
  • 在跨模态问答中表现优于纯文本增强方法
  • 适合需要多模态理解的智能客服与内容生成

多模态检索增强生成(MRAG)通过将文本、图像、视频等多模态数据融入检索与生成过程,提升大语言模型(LLMs)的表现,突破了仅依赖文本的检索增强生成(RAG)局限。尽管传统RAG通过引入外部文本知识提升了回答准确性,但MRAG进一步扩展框架,支持多模态检索与生成,利用多种数据类型的上下文信息,降低生成中的幻觉现象,增强问答系统的事实性。研究表明,MRAG在需同时理解视觉与文本信息的场景中显著优于传统RAG。本文综述了MRAG的核心组件、常用数据集、评估方法及现存局限,分析其构建逻辑与优化路径,并指出关键挑战与未来方向,强调其在多模态信息检索与生成中的变革潜力。该工作为推动该领域发展提供了全面视角。

原文摘要 · Abstract (English)

Multimodal Retrieval-Augmented Generation (MRAG) enhances large language models (LLMs) by integrating multimodal data (text, images, videos) into retrieval and generation processes, overcoming the limitations of text-only Retrieval-Augmented Generation (RAG). While RAG improves response accuracy by incorporating external textual knowledge, MRAG extends this framework to include multimodal retrieval and generation, leveraging contextual information from diverse data types. This approach reduces hallucinations and enhances question-answering systems by grounding responses in factual, multimodal knowledge. Recent studies show MRAG outperforms traditional RAG, especially in scenarios requiring both visual and textual understanding. This survey reviews MRAG's essential components, datasets, evaluation methods, and limitations, providing insights into its construction and improvement. It also identifies challenges and future research directions, highlighting MRAG's potential to revolutionize multimodal information retrieval and generation. By offering a comprehensive perspective, this work encourages further exploration into this promising paradigm.

多模态检索增强大模型生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。