arXiv:2608.27044cs.AIcs.CV2026-08

首个支持文本/视频/音频多模态交互的统一嵌入模型,可按用户任意输入形式检索。

Omni-Interactive Universal Embedder

论文配图:Omni-Interactive Universal Embedder
图 1 · 摘自论文原文
  • 用可学习标记提取中间层特征,构建跨模态统一嵌入空间
  • 在多任务基准上平均提升10.5%至83.7%,尤其在视觉交互任务中提升显著
  • 适合需要多模态交互式检索的研究者与应用开发者

多模态表示学习正从传统双塔架构转向基于大语言模型(LLM)的嵌入器,因其具备强大的指令遵循能力。然而,现有方法主要聚焦于文本和图像模态,且用户交互仍以这两类为主。本文提出首个面向全模态交互的通用嵌入器(OmniUE),不仅通过专用可学习标记的中间层表示,在文本、视频和音频间建立统一嵌入空间,还支持全模态交互查询——用户可用文本、视觉区域或音频片段作为输入。OmniUE中,视觉与音频分割器处理用户交互,并与全模态大语言模型结合,通过上下文聚合生成用户条件化的任意模态间嵌入。为评估其全模态交互能力,我们引入了OmniCHOIR基准,用于基于文本、视频、音频及单模或多模态交互提示的组合音频检索。OmniUE在多种模态下持续超越现有最佳基线,在文本交互视频任务(MMEB-v2-video)上平均提升10.5%,音频任务(MAEB)提升1.1%,视觉交互任务(SCaR)提升83.7%,全模态交互基准(OmniCHOIR)提升24.1%。我们认为,推动全模态表征学习与全模态交互查询的协同发展,将为通用嵌入器铺平道路。

原文摘要 · Abstract (English)

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, existing approaches primarily focus on language and image modalities, which also remain the dominant modalities for user-conditioned interactions in current embedders. In this paper, we propose the first Omni-Interactive Universal Embedder (OmniUE), which not only learns a unified embedding space across text, video, and audio by leveraging intermediate-layer representations from dedicated learnable tokens, but also supports omni-interactive querying, enabling users to provide inputs in the form of text, visual regions of interest, and audio spans. Within OmniUE, visual and audio segmenters process diverse user interactions and integrate them with an omni-LLM to produce user-conditioned any-to-any embeddings via context aggregation. To evaluate OmniUE's omni-interactive capabilities, we introduce OmniCHOIR, benchmarking models for omni-interactive compositional audio retrieval based on the given text, video, and audio as well as unimodal or multimodal interaction prompts. OmniUE consistently surpasses state-of-the-art baselines across diverse modalities, with average improvements of 10.5% on textual-interactive video benchmarks (MMEB-v2-video), 1.1% on audio tasks (MAEB), 83.7% on visual-interactive benchmarks (SCaR), and 24.1% on our omni-interactive OmniCHOIR benchmark. We believe that jointly advancing omni-modal representation learning and omni-interactive querying paves the way toward universal embedders.

多模态交互式检索统一嵌入LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。