arXiv:2505.13880eess.AScs.SD2025-05中稿 · Interspeech 2025被引 11

U-SAM统一理解语音、音频和音乐,通过智能路由与语义感知损失提升跨模态对齐。

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

  • 用专家混合模型动态融合语音、音频、音乐专用编码器输出
  • 在多个基准上超越现有模型,且具备未见任务的涌现能力
  • 适合需要多类型音频统一理解的研究与应用开发者

语音任务的文本生成范式为统一音频理解带来了新可能。然而,现有模型在涵盖语音、通用音频事件和音乐等多样化音频类型时仍面临挑战。此外,其仅依赖交叉熵损失进行对齐,因对所有标记一视同仁且忽略冗余音频特征,导致跨模态对齐效果较弱。为此,本文提出U-SAM,一种集成语音、音频、音乐专用编码器与预训练大语言模型的先进音频语言模型。U-SAM采用专家混合(MoE)投影器实现任务感知的特征融合,动态路由并整合领域特定编码器输出。同时引入语义感知对比损失模块,在语言监督下识别冗余音频特征,并修正其语义与谱表示以增强跨模态对齐。大量实验表明,U-SAM在多个基准上持续优于专用模型与现有音频语言模型,且展现出未见任务上的涌现能力。代码已开源(https://github.com/Honee-W/U-SAM/)。

原文摘要 · Abstract (English)

The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as speech, general audio events, and music. Furthermore, their exclusive reliance on cross-entropy loss for alignment often falls short, as it treats all tokens equally and fails to account for redundant audio features, leading to weaker cross-modal alignment. To deal with the above challenges, this paper introduces U-SAM, an advanced audio language model that integrates specialized encoders for speech, audio, and music with a pre-trained large language model (LLM). U-SAM employs a Mixture of Experts (MoE) projector for task-aware feature fusion, dynamically routing and integrating the domain-specific encoder outputs. Additionally, U-SAM incorporates a Semantic-Aware Contrastive Loss Module, which explicitly identifies redundant audio features under language supervision and rectifies their semantic and spectral representations to enhance cross-modal alignment. Extensive experiments demonstrate that U-SAM consistently outperforms both specialized models and existing audio language models across multiple benchmarks. Moreover, it exhibits emergent capabilities on unseen tasks, showcasing its generalization potential. Code is available (https://github.com/Honee-W/U-SAM/).

音频理解多模态大模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。