视觉AI引入外部知识库,提升理解与生成能力
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook

- 将外部知识检索融入视觉模型,增强上下文感知
- 覆盖图像识别、医疗报告生成、视频3D生成等任务
- 适合研究多模态理解与生成的学者参考
检索增强生成(RAG)已成为人工智能领域关键技术,尤其在提升大语言模型能力方面表现突出,通过接入外部可靠、实时的知识源来增强模型输出。在人工智能生成内容(AIGC)中,RAG能补充相关背景信息,显著提升生成质量。近年来,RAG的应用已拓展至计算机视觉(CV)领域,旨在突破仅依赖内部知识的局限,通过引入权威外部知识库,提升视觉模型的理解与生成能力。本综述系统梳理了当前视觉领域RAG的研究进展,聚焦两大方向:(一)视觉理解,涵盖从基础图像识别到医学报告生成、多模态问答等复杂任务;(二)视觉生成,包括图像、视频与3D内容生成。此外,还探讨了RAG在具身智能中的应用,如规划、任务执行、多模态感知与交互等。鉴于视觉RAG尚处初期阶段,本文亦指出现有方法的主要局限,并提出未来研究方向,以推动该领域的持续发展。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) has emerged as a pivotal technique in artificial intelligence (AI), particularly in enhancing the capabilities of large language models (LLMs) by enabling access to external, reliable, and up-to-date knowledge sources. In the context of AI-Generated Content (AIGC), RAG has proven invaluable by augmenting model outputs with supplementary, relevant information, thus improving their quality. Recently, the potential of RAG has extended beyond natural language processing, with emerging methods integrating retrieval-augmented strategies into the computer vision (CV) domain. These approaches aim to address the limitations of relying solely on internal model knowledge by incorporating authoritative external knowledge bases, thereby improving both the understanding and generation capabilities of vision models. This survey provides a comprehensive review of the current state of retrieval-augmented techniques in CV, focusing on two main areas: (I) visual understanding and (II) visual generation. In the realm of visual understanding, we systematically review tasks ranging from basic image recognition to complex applications such as medical report generation and multimodal question answering. For visual content generation, we examine the application of RAG in tasks related to image, video, and 3D generation. Furthermore, we explore recent advancements in RAG for embodied AI, with a particular focus on applications in planning, task execution, multimodal perception, interaction, and specialized domains. Given that the integration of retrieval-augmented techniques in CV is still in its early stages, we also highlight the key limitations of current approaches and propose future research directions to drive the development of this promising area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。