打造通用音频编码器,融合自监督与有监督模型优势。
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
- 通过领域感知蒸馏解决教师模型不匹配问题。
- 覆盖音乐领域,参数规模达十亿级,性能领先。
- 适合音频大模型研发与多场景音频理解任务。
音频编码器在现代音频应用中至关重要,随着大语言模型(LLM)越来越多地依赖单一编码器处理多样化输入,这一需求愈发突出。尽管自监督学习(SSL)已产出如语音或音乐专家等强领域专用编码器,但多领域方法如USAD和SPEAR仍存在覆盖范围有限、评估不足的问题。近期研究也表明,有监督编码器与音频大模型的对齐效果更优。本文提出USAD 2.0,一种整合自监督与有监督基础模型知识的通用编码器。其引入领域感知蒸馏以缓解教师模型不匹配问题,扩展至音乐领域,并增加第二阶段有监督蒸馏以优化下游应用。通过深度扩展,模型规模扩大至十亿参数。实验表明,USAD 2.0在探针测试与基于大模型的评估中均达到强或顶尖性能。
原文摘要 · Abstract (English)
Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (SSL) has yielded strong domain-specific encoders like speech or music experts, multi-domain approaches like USAD and SPEAR remain limited in coverage and evaluation. Recent studies also suggest supervised encoders align better with audio LLMs. We present USAD 2.0, a universal encoder integrating knowledge from both SSL and supervised foundation models. USAD 2.0 introduces domain-aware distillation to address teacher mismatch, extends coverage to the music domain, and adds second-stage supervised distillation for downstream use. We further scale the model to one billion parameters via depth scaling. Experiments show USAD 2.0 achieves strong or state-of-the-art performance across probing and LLM-based evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。