用语音和音乐模型联合蒸馏,打造更小更强的通用音频模型。
Multi-Distillation from Speech and Music Representation Models
- 用HuBERT和MERT作为教师模型,跨域蒸馏训练统一模型。
- 在多任务上性能媲美专用模型,少样本下表现更优。
- 适合真实场景中数据稀缺的通用音频处理任务。
现实世界音频常同时包含语音和音乐,但现有模型通常只针对单一领域。本文提出一种多教师蒸馏框架,将语音与音乐模型统一为单一模型,显著降低模型规模。方法利用领域专用教师模型(如语音用HuBERT,音乐用MERT)的优势,探索多种平衡策略。在多个任务上的实验表明,该模型性能可媲美专用模型;少样本学习实验进一步凸显通用模型在标注数据有限场景下的重要性。结果证明,跨领域蒸馏不仅有效,且在少样本条件下优于专用模型,验证了跨域统一方法对多样化任务的必要性与优越性。
原文摘要 · Abstract (English)
Real-world audio often mixes speech and music, yet models typically handle only one domain. This paper introduces a multi-teacher distillation framework that unifies speech and music models into a single one while significantly reducing model size. Our approach leverages the strengths of domain-specific teacher models, such as HuBERT for speech and MERT for music, and explores various strategies to balance both domains. Experiments across diverse tasks demonstrate that our model matches the performance of domain-specific models, showing the effectiveness of cross-domain distillation. Additionally, we conduct few-shot learning experiments, highlighting the need for general models in real-world scenarios where labeled data is limited. Our results show that our model not only performs on par with specialized models but also outperforms them in few-shot scenarios, proving that a cross-domain approach is essential and effective for diverse tasks with limited data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。