arXiv:2603.02098cs.IRcs.CL2026-03被引 1

首个支持文本、图像、音频三模态的高效高保真检索模型

Efficient and High-Fidelity Omni Modality Retrieval

  • 用注意力重采样压缩多模态序列,提升计算效率
  • 提出切片沃尔什距离池化,保留细粒度信息提升表征质量
  • 新构建音频中心基准,填补多模态检索评估空白

多模态检索旨在整合跨异构模态的查询信息以检索目标。现有顶尖模型虽能理解复杂查询,但通常仅支持文本与视觉两种模态,限制了通用检索系统的发展。为此,我们提出OmniRet,首个可处理文本、视觉与音频三模态复合查询的检索模型。针对通用检索中的两大挑战——计算效率与表征保真度,我们设计了基于注意力的重采样机制,将各模态编码器输出的海量序列压缩为固定大小表示,显著降低计算开销;同时提出注意力切片沃尔什距离池化(Attention Sliced Wasserstein Pooling),有效保留细粒度特征,增强多模态表征能力。OmniRet在约600万条跨30个数据集的查询-目标对上训练,并在13项检索任务及MMEBv2子集上评估。实验显示,其在复合查询、音频与视频检索任务上表现显著优于基线,其他任务性能与当前最优模型持平。此外,我们构建了新的音频中心多模态基准(ACM),引入复合音频检索与音视频检索两项此前缺失的关键任务,更全面评估模型的全模态嵌入能力。

原文摘要 · Abstract (English)

Multimodal retrieval is the task of aggregating information from queries across heterogeneous modalities to retrieve desired targets. State-of-the-art multimodal retrieval models can understand complex queries, yet they are typically limited to two modalities: text and vision. This limitation impedes the development of universal retrieval systems capable of comprehending queries that combine more than two modalities. To advance toward this goal, we present OmniRet, the first retrieval model capable of handling complex, composed queries spanning three key modalities: text, vision, and audio. Our OmniRet model addresses two critical challenges for universal retrieval: computational efficiency and representation fidelity. First, feeding massive token sequences from modality-specific encoders to Large Language Models (LLMs) is computationally inefficient. We therefore introduce an attention-based resampling mechanism to generate compact, fixed-size representations from these sequences. Second, compressing rich omni-modal data into a single embedding vector inevitably causes information loss and discards fine-grained details. We propose Attention Sliced Wasserstein Pooling to preserve these fine-grained details, leading to improved omni-modal representations. OmniRet is trained on an aggregation of approximately 6 million query-target pairs spanning 30 datasets. We benchmark our model on 13 retrieval tasks and a MMEBv2 subset. Our model demonstrates significant improvements on composed query, audio and video retrieval tasks, while achieving on-par performance with state-of-the-art models on others. Furthermore, we curate a new Audio-Centric Multimodal Benchmark (ACM). This new benchmark introduces two critical, previously missing tasks-composed audio retrieval and audio-visual retrieval to more comprehensively evaluate a model's omni-modal embedding capacity.

多模态检索音频理解高效模型表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。