用极少数据实现尼泊尔语语音克隆,为低资源语言提供个性化合成方案。
Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
- 基于少量数据构建语音编码器与声码器,融合文本与说话人特征生成语音。
- 在未见语音上仍能有效克隆说话人特征,证明小样本语音克隆可行性。
- 为尼泊尔语等低资源语言的个性化语音合成提供可复现的技术基础。
本研究提出一种面向尼泊尔语的少样本语音克隆系统,仅需少量数据即可从天城文文本合成特定说话人的语音。由于尼泊尔语属于低资源语言,相关研究几乎空白。为此,我们分别构建了无标注音频数据集用于训练说话人编码器,以及文本-音频配对数据集用于训练基于Tacotron2的语音合成器。说话人编码器采用生成式端到端损失优化,生成捕捉说话人声纹特征的嵌入向量,通过UMAP降维可视化验证其有效性。这些嵌入与Tacotron2的文本嵌入融合,生成梅尔频谱图,并由WaveRNN声码器还原为语音。音频数据来自自录及其他来源,经严格预处理以保证质量与对齐。训练使用梅尔损失和门控损失,在多种超参数设置下完成。系统在未见语音上仍能有效克隆说话人特征,证实了尼泊尔语少样本语音克隆的可行性,为低资源场景下的个性化语音合成奠定了基础。
原文摘要 · Abstract (English)
This research presents a few-shot voice cloning system for Nepali speakers, designed to synthesize speech in a specific speaker's voice from Devanagari text using minimal data. Voice cloning in Nepali remains largely unexplored due to its low-resource nature. To address this, we constructed separate datasets: untranscribed audio for training a speaker encoder and paired text-audio data for training a Tacotron2-based synthesizer. The speaker encoder, optimized with Generative End2End loss, generates embeddings that capture the speaker's vocal identity, validated through Uniform Manifold Approximation and Projection (UMAP) for dimension reduction visualizations. These embeddings are fused with Tacotron2's text embeddings to produce mel-spectrograms, which are then converted into audio using a WaveRNN vocoder. Audio data were collected from various sources, including self-recordings, and underwent thorough preprocessing for quality and alignment. Training was performed using mel and gate loss functions under multiple hyperparameter settings. The system effectively clones speaker characteristics even for unseen voices, demonstrating the feasibility of few-shot voice cloning for the Nepali language and establishing a foundation for personalized speech synthesis in low-resource scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。