arXiv:2503.22104eess.AS2025-03中稿 · IEEE Access被引 15

融合通用音频与音视频对齐表征,提升跨任务音频理解能力

M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP

  • 联合训练通用音频与音视频对齐特征,利用大模型语义嵌入引导知识迁移
  • 在AudioSet上达到49.0 mAP,音乐任务表现最优
  • 适合需要通用音频表征的多模态应用开发者

对比语言-音频预训练(CLAP)通过将音频与文本对齐于共同特征空间,已成为解决音频任务的流行方法。然而,CLAP的音频特征泛化能力不足,而自监督学习(SSL)模型则能提供跨多样音频任务表现优异的通用特征。本文旨在构建一种广泛适用的音频表征,假设同时学习通用音频特征与CLAP特征可实现目标,即通用音视频表征。为此提出M2D-CLAP,首个联合学习有效通用音频与CLAP特征的方法。它在SSL掩码建模双模型(M2D)基础上引入CLAP,并采用大语言模型(LLM)生成的句子嵌入。训练分多阶段进行:第一阶段通过结合M2D与CLAP的多任务目标预训练通用音频特征,其中CLAP利用LLM语义嵌入将语义知识蒸馏至音频特征;后续阶段在已学音频特征指导下,进一步预训练并优化CLAP特征。实验表明,M2D-CLAP学习到高性能通用音频特征(如AudioSet mAP达49.0,音乐任务取得当前最优结果)及高质量CLAP特征,从而实现通用音视频表征。

原文摘要 · Abstract (English)

Contrastive language-audio pre-training (CLAP), which learns audio-language representations by aligning audio and text in a common feature space, has become popular for solving audio tasks. However, CLAP's audio features lack generalizability, whereas self-supervised learning (SSL) models offer general-purpose features that perform well across diverse audio tasks. We aim to develop a broadly applicable audio representation and hypothesize that a model that learns both general audio and CLAP features should achieve our goal, which we call a general-purpose audio-language representation. To implement our hypothesis, we propose M2D-CLAP, the first approach to jointly learn effective general audio and CLAP features. It extends an SSL masked modeling duo (M2D) by incorporating CLAP and utilizes LLM-based sentence embeddings. The training process consists of multiple stages. In the first stage, generalizable audio features are pre-trained via a multitask objective combining M2D and CLAP, with CLAP leveraging LLM-based semantic embeddings to distill semantic knowledge into them. In the following stages, CLAP features are pre-trained and refined with guidance from the learned audio features. Experiments demonstrated that M2D-CLAP learns high-performing general audio features (e.g., AudioSet mAP of 49.0, SOTA results in music tasks) and CLAP features, thereby enabling a general-purpose audio-language representation.

音视频对齐通用表征自监督学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。