多模态检索增强生成,让大模型更准更懂实时信息。
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
- 用文本、图像、音视频等多模态信息动态补充大模型知识
- 系统梳理了多模态RAG的评估标准与主流技术路径
- 适合关注AI可信生成与跨模态理解的研究者
大型语言模型因依赖静态训练数据,存在幻觉和知识过时问题。检索增强生成(RAG)通过引入外部动态信息提升事实准确性。随着多模态学习的发展,多模态RAG进一步融合文本、图像、音频和视频等多种模态,以增强生成内容的质量。然而,跨模态对齐与推理带来了不同于单模态RAG的独特挑战。本综述对多模态RAG系统进行了系统性分析,涵盖数据集、基准测试、评估指标、方法论创新,以及检索、融合、增强与生成等环节的技术进展。我们还回顾了训练策略、鲁棒性提升、损失函数设计及基于智能体的方法,并探讨了多样化的多模态RAG应用场景。最后,指出了当前开放挑战与未来研究方向,为构建更强大可靠的AI系统提供基础支持。所有资源已公开于 https://github.com/llm-lab-org/Multimodal-RAG-Survey。
原文摘要 · Abstract (English)
Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data. Retrieval-Augmented Generation (RAG) mitigates these issues by integrating external dynamic information for improved factual grounding. With advances in multimodal learning, Multimodal RAG extends this approach by incorporating multiple modalities such as text, images, audio, and video to enhance the generated outputs. However, cross-modal alignment and reasoning introduce unique challenges beyond those in unimodal RAG. This survey offers a structured and comprehensive analysis of Multimodal RAG systems, covering datasets, benchmarks, metrics, evaluation, methodologies, and innovations in retrieval, fusion, augmentation, and generation. We review training strategies, robustness enhancements, loss functions, and agent-based approaches, while also exploring the diverse Multimodal RAG scenarios. In addition, we outline open challenges and future directions to guide research in this evolving field. This survey lays the foundation for developing more capable and reliable AI systems that effectively leverage multimodal dynamic external knowledge bases. All resources are publicly available at https://github.com/llm-lab-org/Multimodal-RAG-Survey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。