arXiv:2508.00955cs.LGcs.AI2025-08中稿 · IEEE Transactions …被引 5

无需大量训练,用提示工程让多模态大模型直接生成高质量嵌入。

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

  • 通过分层提示实现强隐式条件控制,打通模态鸿沟。
  • 仅用少量数据即达基准性能,显著降低训练成本。
  • 适合希望快速部署多模态嵌入的开发者与研究者。

将生成式多模态大模型(MLLMs)转化为通用嵌入模型通常需要资源密集型的对比预训练,而传统硬负样本挖掘方法易受错误负样本污染。本文提出一种高数据效率框架,无需大规模预训练即可构建鲁棒的多模态表示空间。首先引入分层嵌入提示,从系统层面显式锚定任务定义,有效弥合模态差异,释放强大的零样本嵌入能力。在此基础上,提出自感知硬负样本采样(SaHa),不依赖候选集挖掘,而是将检索结果映射回原始查询,严格过滤语义错误负例。此外,该方法构建相互困难簇,在不增加前向传播的前提下最大化任务内区分度与批量效率。大量实验表明,本方法在仅使用标准训练数据一小部分的情况下,于大规模多模态嵌入基准上达到具有竞争力的微调性能。

原文摘要 · Abstract (English)

Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative contamination. In this paper, we propose a highly data-efficient framework that bypasses extensive pre-training to build a robust multimodal representation space. We first introduce a hierarchical embedding prompt that provides strong latent conditioning. By explicitly anchoring task definitions at the system level, this prompting strategy effectively bridges the modality gap and unlocks powerful zero-shot embedding capabilities. Building upon this latent conditioning, we present Self-aware Hard Negative Sampling (SaHa). Unlike conventional candidate-space mining, SaHa shifts the mechanism to the query-space by mapping retrieved candidates back to their owner queries to rigorously filter out semantic false negatives. Furthermore, our method constructs mutually hard clusters, maximizing intra-task discrimination and batch efficiency without redundant forward passes. Extensive experiments demonstrate that our unified approach achieves highly competitive fine-tuning performance on the Massive Multimodal Embedding Benchmark using only a fraction of standard training data.

多模态嵌入提示工程零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。