用语音增强模型迁移训练唱歌分离模型,数据少也能高效提升效果。
Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation
- 从语音增强模型迁移,通过微调实现唱歌分离
- 性能提升0.29-1.8 dB,LoRA仅增6-12%参数
- 适合缺乏唱歌数据的场景,尤其适合资源有限的研究者
当前先进的语音增强模型依赖大规模标注数据集,而歌唱语音分离模型则受限于可用训练数据不足。为解决此问题,本文将歌唱语音分离建模为从语音增强到歌唱语音分离的领域自适应任务。研究了两种微调策略:全量微调与基于低秩适配(LoRA)的参数高效微调,分别应用于判别式和生成式模型。采用任一适应策略的模型均在信噪比(SDR)上优于从零训练的同架构模型,提升0.29至1.8 dB。全量微调表现最佳,但导致语音增强能力严重退化;而LoRA微调在仅增加6%-12%参数的前提下,保持原始语音增强性能的同时获得有竞争力的歌唱分离效果。此外,生成式模型在未见测试集上表现出更强泛化能力。结果表明,在数据稀缺情况下,迁移预训练语音增强模型是训练歌唱语音分离模型的有效策略。
原文摘要 · Abstract (English)
State-of-the-art speech enhancement models benefit from large-scale labeled datasets, whereas singing voice separation models suffer from limited available training data. To address this limitation, we formulate singing voice separation as domain adaptation from speech enhancement to singing voice separation. We investigate two fine-tuning strategies: full fine-tuning and parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) on a discriminative and a generative model. Models with either adaptation strategy outperform the same architectures trained from scratch by 0.29-1.8 dB in Signal-to-Distortion-Ratio. Full fine-tuning yields the highest singing voice separation performance, but catastrophic forgetting degrades speech enhancement performance. LoRA fine-tuning achieves competitive singing voice separation performance while preserving the original speech enhancement capability with only 6-12% additional parameters compared to the base speech enhancement model. Furthermore, the generative model shows improved generalization to an unseen test set. The results demonstrate that adapting pretrained speech enhancement models is an effective strategy for training singing voice separation models in data-scarce scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。