arXiv:2410.14154cs.MMcs.AI2024-10被引 11

用自适应检索增强训练,让多模态模型更懂图中文义。

RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training

  • 用问题引导视觉信息提取,减少无关干扰。
  • 融合图文到统一语义空间,提升知识检索准确率。
  • 自动筛选相关知识,适合需要实时更新的多模态应用。

多模态大语言模型(MLLMs)在视觉-语言任务中展现出强大潜力,但其内部存储的外部知识难以持续更新,存在计算成本高、可解释性差的问题。检索增强技术已被证明对语言模型和多模态模型有效。本文提出一种新的多模态自适应检索增强自举预训练框架(RA-BLIP)。针对视觉模态中的冗余信息,利用问题指导一组可学习查询,交互式提取关键视觉信息,降低检索与生成过程中的无关干扰。此外,引入预训练的多模态自适应融合模块,将视觉与语言模态投影至统一语义空间,实现问题文本到多模态知识的检索与整合。进一步设计自适应选择知识生成(ASKG)策略,使生成器自主判断检索内容的相关性,实现优异的去噪性能。在多个开放多模态问答数据集上的实验表明,RA-BLIP性能显著优于现有检索增强模型。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have recently received substantial interest, which shows their emerging potential as general-purpose models for various vision-language tasks. MLLMs involve significant external knowledge within their parameters; however, it is challenging to continually update these models with the latest knowledge, which involves huge computational costs and poor interpretability. Retrieval augmentation techniques have proven to be effective plugins for both LLMs and MLLMs. In this study, we propose multimodal adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training (RA-BLIP), a novel retrieval-augmented framework for various MLLMs. Considering the redundant information within vision modality, we first leverage the question to instruct the extraction of visual information through interactions with one set of learnable queries, minimizing irrelevant interference during retrieval and generation. Besides, we introduce a pre-trained multimodal adaptive fusion module to achieve question text-to-multimodal retrieval and integration of multimodal knowledge by projecting visual and language modalities into a unified semantic space. Furthermore, we present an Adaptive Selection Knowledge Generation (ASKG) strategy to train the generator to autonomously discern the relevance of retrieved knowledge, which realizes excellent denoising performance. Extensive experiments on open multimodal question-answering datasets demonstrate that RA-BLIP achieves significant performance and surpasses the state-of-the-art retrieval-augmented models.

多模态检索增强自适应知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。