让图文嵌入模型能通过点击/框选区域来理解用户意图,提升定位精准度。
VIRTUE: Visual-Interactive Text-Image Universal Embedder
- 引入视觉交互机制,支持点选、框选等区域提示输入。
- 在36项通用任务上提升3.1%-8.5%,在5项交互任务上提升15.2%-20.3%。
- 适合需要精准图像语义定位的多模态应用开发人员。
多模态表示学习模型已在复杂任务中表现优异,视觉语言模型(VLM)的融合进一步使嵌入模型具备指令遵循能力。然而,现有嵌入模型缺乏视觉交互能力,无法响应用户对图像特定区域的指定(如点选、边界框、掩码),而这类能力在生成模型中已被探索以增强人机交互性。赋予嵌入模型视觉交互能力不仅能开启基于局部语义定位的新应用,还能帮助模型学习图像中的实体级信息,补充其全局表征。本文提出一种新型视觉-交互式文本-图像通用嵌入模型(VIRTUE),将分割模型与视觉语言模型的能力延伸至表示学习领域。在VIRTUE中,分割模型可处理指向图像特定区域的视觉提示,从而更精确应对复杂模糊场景。为评估其视觉交互能力,我们构建了一个包含100万样本的大规模分割与场景描述检索(SCaR)基准,旨在通过联合考虑特定对象实体与图像场景来检索文本描述。VIRTUE在36项通用多模态嵌入基准(MMEB)任务上均达到领先水平,性能提升3.1%-8.5%;在5项视觉交互型SCaR任务上提升15.2%-20.3%。
原文摘要 · Abstract (English)
Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embedding models with instruction-following capabilities. However, existing embedding models lack visual-interactive capabilities to specify regions of interest from users (e.g., point, bounding box, mask), which have been explored in generative models to broaden their human-interactive applicability. Equipping embedding models with visual interactions not only would unlock new applications with localized grounding of user intent, which remains unexplored, but also enable the models to learn entity-level information within images to complement their global representations for conventional embedding tasks. In this paper, we propose a novel Visual-InteRactive Text-Image Universal Embedder (VIRTUE) that extends the capabilities of the segmentation model and the vision-language model to the realm of representation learning. In VIRTUE, the segmentation model can process visual prompts that pinpoint specific regions within an image, thereby enabling the embedder to handle complex and ambiguous scenarios more precisely. To evaluate the visual-interaction ability of VIRTUE, we introduce a large-scale Segmentation-and-Scene Caption Retrieval (SCaR) benchmark comprising 1M samples that aims to retrieve the text caption by jointly considering the entity with a specific object and image scene. VIRTUE consistently achieves a state-of-the-art performance with significant improvements across 36 universal MMEB (3.1%-8.5%) and five visual-interactive SCaR (15.2%-20.3%) tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。