用蒸馏统一学习语音与音频,一个模型搞定多种任务。
USAD: Universal Speech and Audio Representation via Distillation
- 从多个专用模型蒸馏知识,训练统一的音频表示模型。
- 在SUPERB和HEAR基准上达到接近顶尖水平的表现。
- 适合需要多类型音频处理的开发者和研究者使用。
自监督学习(SSL)已革新音频表征,但现有模型多局限于特定领域,专注语音或非语音任务。本文提出通用语音与音频蒸馏(USAD),一种统一的音频表征学习方法,将语音、声音和音乐等多种音频类型整合至单一模型。USAD通过从领域专用的SSL模型中进行高效层间蒸馏,在大规模音频数据集上训练学生模型。该方法在多个基准测试中表现优异,涵盖帧级与实例级语音处理、音频标注及声音分类等任务,在SUPERB和HEAR基准上以单一编码器实现接近当前最优的结果。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types - speech, sound, and music - into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。