构建多模态嵌入评估新基准,揭示当前模型在跨模态检索中的关键缺陷。
MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

- 提出MMEB-V3与OmniSET,实现全模态统一评测与语义对齐诊断。
- 发现模型跨模态检索存在严重不对称性,且指令引导效果不佳。
- 适合研究多模态对齐、评估体系与指令理解的学者参考。
多模态嵌入模型旨在将文本、图像、视频、音频等异构输入映射到共享语义空间。然而,现有方法与基准大多仅覆盖部分模态,难以系统评估全模态表征学习。本文迈向全模态设定,提出MMEB-V3——一个涵盖文本、图像、视频、音频及以代理为中心场景的综合性评测基准。为实现更细粒度诊断,我们构建OmniSET(全模态语义等价对),使跨模态的语义等价实例得以表示,从而分离语义相似性与模态效应。在MMEB-V3上的实验揭示三大发现:(1) 模型常无法检索目标模态;(2) 跨模态检索高度不对称,受查询模态偏见主导;(3) 指令诱导的语义迁移或不足或与目标模态错位,无法可靠提升检索性能。结果表明,当前多模态嵌入尚无法可靠执行指令指定的模态约束,因而缺乏一致的模态感知检索行为。我们期望MMEB-V3能为理解与诊断这些局限提供有力工具,并指导未来全模态嵌入研究。
原文摘要 · Abstract (English)
Multimodal embedding models aim to map heterogeneous inputs, such as text, images, videos, and audio, into a shared semantic space. However, existing methods and benchmarks remain largely limited to partial modality coverage, making it difficult to systematically evaluate full-modality representation learning. In this work, we take a step toward the full-modality setting. We introduce MMEB-V3, a comprehensive benchmark that evaluates embeddings across text, image, video, audio, as well as agent-centric scenarios. To enable more fine-grained diagnosis, we further construct OmniSET (Omni-modality Semantic Equivalence Tuples), where semantically equivalent instances are represented across modalities, allowing us to disentangle semantic similarity from modality effects. Through experiments on MMEB-V3, we conduct a systematic analysis of full-modality embeddings and identify three key findings: (1) models often fail to retrieve the intended target modality; (2) cross-modal retrieval is highly asymmetric and dominated by query-modality bias; and (3) instruction-induced shifts are either insufficient or misaligned with the target modality, and therefore do not reliably improve retrieval. These results indicate that current multimodal embeddings are not yet capable of reliably enforcing modality constraints specified by instructions, and consequently fail to exhibit consistent modality-aware retrieval behavior. We hope MMEB-V3 provides a useful benchmark for understanding and diagnosing these limitations, and for guiding future research on full-modality embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。