arXiv:2412.16855cs.CLcs.IR2024-12中稿 · CVPR被引 188

用合成数据提升多模态检索模型性能

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

论文配图:GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
图 1 · 摘自论文原文
  • 基于多模态大模型构建融合模态训练数据
  • 在多个基准上达到当前最优检索效果
  • 适合研究多模态检索与大模型应用者

通用多模态检索(UMR)旨在通过统一模型实现跨模态搜索,查询和候选可为纯文本、图像或两者组合。以往工作尝试仅用文本数据利用多模态大语言模型(MLLMs)实现UMR,但实验表明更丰富的多模态训练数据能进一步释放MLLM潜力。现有数据在模态分布上高度不均衡,为此我们设计了数据合成管道,构建大规模高质量的融合模态训练集。基于此,提出通用多模态嵌入器(GME),一种基于MLLM的密集检索模型。此外,构建了全面的UMR基准(UMRB)评估方法有效性。实验结果表明,该方法在现有UMR方法中表现最佳。最后,对模型扩展性、训练策略进行了深入分析,并对模型与合成数据进行消融研究。

原文摘要 · Abstract (English)

Universal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt multimodal large language models (MLLMs) to realize UMR using only text data. However, our preliminary experiments demonstrate that more diverse multimodal training data can further unlock the potential of MLLMs. Despite its effectiveness, the existing multimodal training data is highly imbalanced in terms of modality, which motivates us to develop a training data synthesis pipeline and construct a large-scale, high-quality fused-modal training dataset. Based on the synthetic training data, we develop the General Multimodal Embedder (GME), an MLLM-based dense retriever designed for UMR. Furthermore, we construct a comprehensive UMR Benchmark (UMRB) to evaluate the effectiveness of our approach. Experimental results show that our method achieves state-of-the-art performance among existing UMR methods. Last, we provide in-depth analyses of model scaling and training strategies, and perform ablation studies on both the model and synthetic data.

多模态检索大模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。