零样本跨语言歌声转换,支持多语言无缝切换。
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
- 用SPIN增强VITS模型提取通用语音内容特征
- 采用ECAPA2说话人编码器分离音色与语言信息
- 无需特定语言训练,适合多语言歌声合成场景
本文提出FreeSVC,一种基于改进VITS模型的多语言歌声转换方法,结合说话人无关聚类(SPIN)以提升内容表征能力,并采用当前最优的说话人编码器ECAPA2。FreeSVC引入可训练的语言嵌入,实现多语言支持,通过先进编码器将说话人特征与语言内容解耦。该方法专为零样本学习设计,可在无特定语言训练数据的情况下实现跨语言歌声转换。实验表明,多语言内容提取器对跨语言转换性能至关重要。项目代码与模型已公开。
原文摘要 · Abstract (English)
This work presents FreeSVC, a promising multilingual singing voice conversion approach that leverages an enhanced VITS model with Speaker-invariant Clustering (SPIN) for better content representation and the State-of-the-Art (SOTA) speaker encoder ECAPA2. FreeSVC incorporates trainable language embeddings to handle multiple languages and employs an advanced speaker encoder to disentangle speaker characteristics from linguistic content. Designed for zero-shot learning, FreeSVC enables cross-lingual singing voice conversion without extensive language-specific training. We demonstrate that a multilingual content extractor is crucial for optimal cross-language conversion. Our source code and models are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。