arXiv:2606.30682cs.SDcs.AI2026-06

用大模型学通用音频嵌入,支持指令控制的多任务检索

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

论文配图:ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models
图 1 · 摘自论文原文
  • 从大音视频模型迁移能力,构建统一音频嵌入空间
  • 在标准数据集上表现优异,支持组合式和可控检索
  • 适合需要灵活指令响应的语音检索场景

近期的音视频检索进展主要依赖对比双编码器架构,在共享嵌入空间中对齐音频与文本。尽管有效,现有嵌入模型主要针对音频-字幕匹配优化,难以支持多样化检索目标和可控检索行为。本文提出ALM2Vec,一种源自预训练大型音视频模型(LALMs)的通用音频嵌入框架。通过迁移大规模多模态训练中获得的音频理解、指令遵循与推理能力,ALM2Vec学习一个跨音频领域与任务类型的统一嵌入空间。除传统文本-音频检索外,还引入自然语言指令,实现面向音频问答与属性条件检索等指令感知检索。实验表明,ALM2Vec在标准音视频与语音检索基准上表现优异,并展现出出色的组合性与可控性检索能力,凸显其作为跨领域、跨任务、跨用户意图的统一音频嵌入模型的潜力。

原文摘要 · Abstract (English)

Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effective, existing retrieval embeddings are primarily optimized for audio--caption matching, limiting their ability to support diverse retrieval objectives and controllable retrieval behaviors. We present ALM2Vec, a universal audio embedding framework derived from pretrained large audio--language models (LALMs). By transferring the audio understanding, instruction-following, and reasoning capabilities acquired through large-scale multimodal training, ALM2Vec learns a unified embedding space for retrieval across audio domains and task types. Beyond conventional text--audio retrieval, ALM2Vec incorporates natural-language instructions into the embedding process, enabling instruction-aware retrieval for scenarios such as audio question answering and aspect-conditioned retrieval. Experimental results show that ALM2Vec achieves competitive performance on standard audio and speech retrieval benchmarks while exhibiting promising compositional and controllable retrieval capabilities, highlighting its potential as a unified audio embedding model for retrieval across domains, tasks, and user intents.

音频检索大模型指令控制嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。