提出音频文本交错检索任务,提升多模态信息查找能力。
ATIR: Towards Audio-Text Interleaved Contextual Retrieval

- 设计音频与文本交替的检索任务,统一四类上下文检索场景。
- 构建ATIR基准,融合多个语音识别与问答数据集,支持语义级检索。
- 引入新令牌压缩机制,解决大模型中音频令牌过量问题。
音频蕴含比文本更丰富的信息,包括情感、说话人特征和环境背景,且相比语音转文字流程具有更低延迟。然而,当前多模态信息检索研究主要聚焦图像,忽视了音频,尤其在音频与文本交错的上下文检索场景中。本文提出音频-文本交错上下文检索(ATIR)任务,允许查询在音频与文本模态间交替。我们通过整合多个自动语音识别(ASR)、问答(QA)和检索数据集,构建了ATIR基准,最终统一了四类上下文检索任务,显著弥补了现有音频检索数据集在语义检索方面的不足。为研究该任务,我们评估了多种现成检索器,并基于多模态大语言模型(MLLM)训练了ATIR模型。进一步提出一种与现有压缩方法正交的新令牌压缩机制,有效缓解了基于MLLM的ATIR模型中音频令牌过多的问题。实验表明,所提ATIR模型显著优于强基线。
原文摘要 · Abstract (English)
Audio carries richer information than text, including emotion, speaker traits, and environmental context, while also enabling lower-latency processing compared to speech-to-text pipelines. However, recent multimodal information retrieval research has predominantly focused on images, largely overlooking audio, especially in the setting of interleaved audio-text contextual retrieval. In this work, we introduce the Audio-Text Interleaved contextual Retrieval (ATIR) task, where queries can alternate between audio and text modalities. We construct an ATIR benchmark by integrating several Automatic Speech Recognition (ASR), QA, and retrieval datasets, ultimately unifying four types of contextual retrieval tasks. This benchmark substantially addresses the limitations of existing audio retrieval datasets in semantic retrieval. To study this task, we evaluate several off-the-shelf retrievers and train our ATIR model based on a Multimodal Large Language Model (MLLM). We further introduce a novel token compression mechanism that is orthogonal to existing compression methods, thereby alleviating the issue of excessive audio tokens in MLLM-based ATIR models. Experimental results demonstrate that our ATIR model achieves substantial improvements over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。