arXiv:2506.18843cs.SDcs.CL2025-06中稿 · ASRU 2025被引 12

用蒸馏统一学习语音与音频,一个模型搞定多种任务。

USAD: Universal Speech and Audio Representation via Distillation

  • 从多个专用模型蒸馏知识,训练统一的音频表示模型。
  • 在SUPERB和HEAR基准上达到接近顶尖水平的表现。
  • 适合需要多类型音频处理的开发者和研究者使用。

自监督学习(SSL)已革新音频表征,但现有模型多局限于特定领域,专注语音或非语音任务。本文提出通用语音与音频蒸馏(USAD),一种统一的音频表征学习方法,将语音、声音和音乐等多种音频类型整合至单一模型。USAD通过从领域专用的SSL模型中进行高效层间蒸馏,在大规模音频数据集上训练学生模型。该方法在多个基准测试中表现优异,涵盖帧级与实例级语音处理、音频标注及声音分类等任务,在SUPERB和HEAR基准上以单一编码器实现接近当前最优的结果。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks.

音频表征自监督学习蒸馏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。