arXiv:2412.01720cs.CV2024-12CVPR被引 117

用大模型当检索助手,一套方法搞定多种检索任务

LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

论文配图:LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
图 1 · 摘自论文原文
  • 让大模型通过两阶段训练掌握检索能力
  • 在十余项任务上表现优异,零样本也能用
  • 适合需要灵活应对新检索需求的研究者

随着多模态信息检索的快速发展,越来越多复杂的检索任务出现。现有方法主要依赖针对特定任务微调的视觉-语言模型,通常基于图像-文本对比学习训练。本文探索将生成式大型多模态模型(LMMs)重新用于检索任务。该方法可统一所有检索任务的处理框架,并能无需额外训练即泛化到未见过的检索任务。我们的贡献包括:(i) 提出 LamRA 框架,赋予 LMM 强大的检索与重排序能力;(ii) 采用仅语言预训练+多模态指令微调的两阶段策略,逐步提升检索性能;(iii) 重排序采用点对点与列表级联合训练,提供两种增强方式;(iv) 大量实验表明,该方法在超过十种检索任务中表现优异,无论监督或零样本场景均有效,包括从未见过的任务。

原文摘要 · Abstract (English)

With the rapid advancement of multimodal information retrieval, increasingly complex retrieval tasks have emerged. Existing methods predominately rely on task-specific fine-tuning of vision-language models, often those trained with image-text contrastive learning. In this paper, we explore the possibility of re-purposing generative Large Multimodal Models (LMMs) for retrieval. This approach enables unifying all retrieval tasks under the same formulation and, more importantly, allows for extrapolation towards unseen retrieval tasks without additional training. Our contributions can be summarised in the following aspects: (i) We introduce LamRA, a versatile framework designed to empower LMMs with sophisticated retrieval and reranking capabilities. (ii) For retrieval, we adopt a two-stage training strategy comprising language-only pre-training and multimodal instruction tuning to progressively enhance LMM's retrieval performance. (iii) For reranking, we employ joint training for both pointwise and listwise reranking, offering two distinct ways to further boost the retrieval performance. (iv) Extensive experimental results underscore the efficacy of our method in handling more than ten retrieval tasks, demonstrating robust performance in both supervised and zero-shot settings, including scenarios involving previously unseen retrieval tasks.

多模态检索大模型应用零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。