arXiv:2503.04812cs.CVcs.AI2025-03EMNLP被引 60

用难负样本加权对比学习,提升图文嵌入模型的区分能力

LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning

  • 基于判别难度动态调整负样本权重,优化嵌入表示
  • LLaVE-7B在36个数据集上比前序SOTA高6.2分
  • 零样本迁移至文-视频检索,展现强泛化潜力

通用多模态嵌入模型在图文混排检索、多模态RAG和聚类等任务中至关重要。然而,我们发现现有基于LMM的嵌入模型使用标准InfoNCE损失训练时,正负样本相似度分布重叠严重,难以有效区分难负样本。为此,我们提出一种简单而有效的框架,根据负样本的判别难度动态增强其对嵌入学习的贡献。在此框架下,我们训练了一系列名为LLaVE的模型,并在涵盖4个元任务和36个数据集的MMEB基准上进行评估。实验表明,LLaVE建立了更强基线,在保持高效性和可扩展性的同时达到当前最优(SOTA)性能:LLaVE-2B超越此前7B规模的SOTA模型,而LLaVE-7B进一步提升6.2分。尽管仅在图文数据上训练,LLaVE仍能以零样本方式迁移到文-视频检索任务并取得优异表现,展现出向其他嵌入任务迁移的强大潜力。

原文摘要 · Abstract (English)

Universal multimodal embedding models play a critical role in tasks such as interleaved image-text retrieval, multimodal RAG, and multimodal clustering. However, our empirical results indicate that existing LMM-based embedding models trained with the standard InfoNCE loss exhibit a high degree of overlap in similarity distribution between positive and negative pairs, making it challenging to distinguish hard negative pairs effectively. To deal with this issue, we propose a simple yet effective framework that dynamically improves the embedding model's representation learning for negative pairs based on their discriminative difficulty. Within this framework, we train a series of models, named LLaVE, and evaluate them on the MMEB benchmark, which covers 4 meta-tasks and 36 datasets. Experimental results show that LLaVE establishes stronger baselines that achieve state-of-the-art (SOTA) performance while demonstrating strong scalability and efficiency. Specifically, LLaVE-2B surpasses the previous SOTA 7B models, while LLaVE-7B achieves a further performance improvement of 6.2 points. Although LLaVE is trained on image-text data, it can generalize to text-video retrieval tasks in a zero-shot manner and achieve strong performance, demonstrating its remarkable potential for transfer to other embedding tasks.

多模态嵌入对比学习零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。