NanoVoice高效实现多说话人语音合成,训练快4倍、参数少45%
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
- 批量并行微调多个参考语音,实现多说话人自适应
- 40个说话人时训练速度提升4倍,参数减少45%
- 适合需要快速部署多说话人语音系统的场景
我们提出NanoVoice,一种高效的个性化文本到语音模型,可同时为多个说话人构建语音适配器。该模型引入批处理式说话人适配技术,能并行微调多个参考语音,显著缩短训练时间。除为每位说话人分别构建适配器外,还提出参数共享机制,减少说话人适配所需参数量。通过引入可学习的缩放矩阵,有效缓解参数共享带来的性能下降。在40个参考语音下,NanoVoice性能媲美基线模型,训练速度提升4倍,适配参数减少45%。大量消融实验和分析进一步验证了模型效率。
原文摘要 · Abstract (English)
We present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker adaptation technique capable of fine-tuning multiple references in parallel, significantly reducing training time. Beyond building separate adapters for each speaker, we also propose a parameter sharing technique that reduces the number of parameters used for speaker adaptation. By incorporating a novel trainable scale matrix, NanoVoice mitigates potential performance degradation during parameter sharing. NanoVoice achieves performance comparable to the baselines, while training 4 times faster and using 45 percent fewer parameters for speaker adaptation with 40 reference voices. Extensive ablation studies and analysis further validate the efficiency of our model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。