arXiv:2504.07858cs.SDcs.AI2025-04

用少量数据实现高质量语音合成,让小语种也能有自然声音。

Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis

  • 结合优化数据与先进声学模型,提升低资源语言合成效果。
  • 在泰语上实现零样本语音克隆,性能达当前最佳水平。
  • 适合金融、医疗等需要多语言语音服务的场景使用。

文本转语音(TTS)技术在主流语言上已取得显著进展,但许多资源匮乏的语言仍面临数据稀缺与语言复杂性的挑战。本文提出一种新型方法,将数据优化框架与先进声学模型相结合,构建适用于低资源场景的高质量TTS系统。以泰语为例,该方法有效应对复杂的语音规则和稀疏数据问题。实验表明,该方法支持零样本语音克隆,并在金融、医疗、教育、法律等多个客户端应用中表现优异。主观与客观评估均证实,模型达到当前最先进水平,为数据受限环境下的TTS规模化生产提供了可行方案,对推动多语言可及性与行业应用具有重要意义。

原文摘要 · Abstract (English)

Text-to-speech (TTS) technology has achieved impressive results for widely spoken languages, yet many under-resourced languages remain challenged by limited data and linguistic complexities. In this paper, we present a novel methodology that integrates a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios. We demonstrate the effectiveness of our approach using Thai as an illustrative case, where intricate phonetic rules and sparse resources are effectively addressed. Our method enables zero-shot voice cloning and improved performance across diverse client applications, ranging from finance to healthcare, education, and law. Extensive evaluations - both subjective and objective - confirm that our model meets state-of-the-art standards, offering a scalable solution for TTS production in data-limited settings, with significant implications for broader industry adoption and multilingual accessibility.

语音合成低资源语言零样本克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。