arXiv:2507.04590cs.CVcs.CL2025-07被引 107

统一建模视频、图像与图文文档的嵌入表示,提升多模态应用泛化能力。

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

  • 构建跨视觉形式的统一嵌入框架,支持文本、图像、视频和文档输入。
  • 在新增的视频检索与文档检索任务上超越基线,同时提升原有图像基准表现。
  • 适用于AI代理、多模态搜索与RAG等实际场景,推动通用视觉表征发展。

多模态嵌入模型在语义相似性、信息检索和聚类等下游任务中至关重要。然而,现有模型如VLM2Vec、E5-V、GME主要聚焦自然图像,对视频和视觉文档的支持有限,限制了其在真实场景中的应用,包括AI代理、多模态搜索与推荐、检索增强生成(RAG)。为填补这一空白,我们提出VLM2Vec-V2,一个面向多样化视觉形式的统一嵌入框架。首先,我们引入MMEB-V2,该基准扩展了MMEB,包含五种新任务类型:视觉文档检索、视频检索、时序定位、视频分类和视频问答,覆盖文本、图像、视频与视觉文档输入。随后,我们训练了VLM2Vec-V2,一个支持文本、图像、视频和视觉文档输入的通用嵌入模型。大量实验表明,VLM2Vec-V2不仅在新引入的视频与文档检索任务上表现优异,还在原有图像基准上优于先前基线。通过全面评估,本研究揭示了多种多模态嵌入模型的泛化能力,并指出有效的统一嵌入学习策略,为研究与实际应用中的可扩展、自适应表征学习奠定基础。

原文摘要 · Abstract (English)

Multimodal embedding models have been crucial in enabling various downstream tasks such as semantic similarity, information retrieval, and clustering over different modalities. However, existing multimodal embeddings like VLM2Vec, E5-V, GME are predominantly focused on natural images, with limited support for other visual forms such as videos and visual documents. This restricts their applicability in real-world scenarios, including AI agents, multi-modal search and recommendation, and retrieval-augmented generation (RAG). To close this gap, we propose VLM2Vec-V2, a unified framework for learning embeddings across diverse visual forms. First, we introduce MMEB-V2, a comprehensive benchmark that extends MMEB with five new task types: visual document retrieval, video retrieval, temporal grounding, video classification and video question answering - spanning text, image, video, and visual document inputs. Next, we train VLM2Vec-V2, a general-purpose embedding model that supports text, image, video, and visual document inputs. Extensive experiments show that VLM2Vec-V2 achieves strong performance not only on the newly introduced video and document retrieval tasks, but also improves over prior baselines on the original image benchmarks. Through extensive evaluation, our study offers insights into the generalizability of various multimodal embedding models and highlights effective strategies for unified embedding learning, laying the groundwork for more scalable and adaptable representation learning in both research and real-world settings.

多模态嵌入视频理解文档检索统一表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。