首次系统评估扩散模型作多模态嵌入,发现其表现普遍弱于自回归模型。
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
- 将扩散模型转换为嵌入模型,对比其在三类任务中的表现。
- LaViDa在分类、VQA和检索上分别落后3.5、2.5、4.4分。
- 扩散模型图像-文本对齐不足是性能短板,适合研究多模态对齐的学者。
嵌入模型是现代AI系统(如语义搜索与检索增强生成)的核心组件。近年来,大基础模型的进展显著加速了嵌入模型的发展,包括基于大语言模型(LLMs)、视觉语言模型(VLMs)及多模态LLMs的模型。最近,大扩散语言模型(dLLMs)和多模态dLLMs作为自回归模型的有力替代者出现,具备双向注意力和并行生成等优势。这引出一个关键但未被探索的问题:多模态dLLMs能否作为有效的多模态嵌入模型?为此,我们首次系统性地研究将多模态dLLMs转化为嵌入模型的方法。我们在分类、视觉问答(VQA)和信息检索三类嵌入任务上,评估了当前先进的多模态dLLMs与自回归VLMs。结果表明,多模态dLLM嵌入在多数情况下逊于自回归VLM。其中,更强的扩散模型LaViDa在分类、VQA和检索任务上分别落后3.5、2.5、4.4分;而另一扩散模型MMaDA则在所有任务上差距超过20分。进一步分析显示,扩散模型中图像-文本对齐不足是其嵌入性能受限的主要原因。
原文摘要 · Abstract (English)
Embedding models are a fundamental component of modern AI systems such as semantic search and retrieval-augmented generation. Recent advances in large foundation models have substantially accelerated the development of embedding models, including those based on Large Language Models (LLMs), Vision Language Models (VLMs), and Multimodal LLMs. More recently, Large Diffusion Language Models (dLLMs) and Multimodal dLLMs have emerged as competitive alternatives to autoregressive models, offering advantages such as bidirectional attention and parallel generation. This progress naturally raises a critical yet unexplored question: can Multimodal dLLMs serve as effective multimodal embedding models? To answer this, we present the first systematic study of converting Multimodal dLLMs into embedding models. We evaluate state-of-the-art Multimodal dLLMs and Autoregressive VLMs across three categories of embedding tasks: classification, visual question answering, and information retrieval. Our results show that Multimodal dLLM embeddings generally underperform their autoregressive VLM counterparts. The stronger diffusion-based model, LaViDa, lags by only 3.5 points on classification, 2.5 points on VQA, and 4.4 points on retrieval tasks, whereas the other diffusion-based model, MMaDA, exhibits substantially larger performance gaps, exceeding 20 points across all tasks. Further analysis reveals insufficient image-text alignment in diffusion-based models, accounting for the observed limitations in their embedding performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。