打造支持六种客家话的高自然度语音合成系统,助力濒危语言保护。
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
- 基于YourTTS框架,融合方言特异性数据训练,实现多方言多说话人合成
- 通过网页抓取与ASR清洗,构建高质量多说话人多方言语料库
- 主观评测显示在发音准确、声调正确和整体自然度上显著优于现有系统
本文介绍VoxHakka,一个专为台湾客家话设计的文本转语音(TTS)系统。该语言资源匮乏,而VoxHakka基于YourTTS框架,实现了高自然度、高准确率和低实时因子的语音合成,并支持六种不同客家话方言。通过使用方言特异性数据训练模型,实现说话人感知的客家话语音生成。为解决公开可用的客家话语音语料库稀缺问题,研究采用低成本的网页抓取管道结合自动语音识别(ASR)数据清洗技术,获取高质量、多说话人、多方言的数据集,适用于TTS训练。主观听感测试采用对比平均意见分(CMOS)评估,结果表明VoxHakka在发音准确性、声调正确性和整体自然度方面显著优于现有公开的客家话TTS系统。本工作推动了客家话语言技术发展,为语言保护与复兴提供了宝贵资源。
原文摘要 · Abstract (English)
This paper introduces VoxHakka, a text-to-speech (TTS) system designed for Taiwanese Hakka, a critically under-resourced language spoken in Taiwan. Leveraging the YourTTS framework, VoxHakka achieves high naturalness and accuracy and low real-time factor in speech synthesis while supporting six distinct Hakka dialects. This is achieved by training the model with dialect-specific data, allowing for the generation of speaker-aware Hakka speech. To address the scarcity of publicly available Hakka speech corpora, we employed a cost-effective approach utilizing a web scraping pipeline coupled with automatic speech recognition (ASR)-based data cleaning techniques. This process ensured the acquisition of a high-quality, multi-speaker, multi-dialect dataset suitable for TTS training. Subjective listening tests conducted using comparative mean opinion scores (CMOS) demonstrate that VoxHakka significantly outperforms existing publicly available Hakka TTS systems in terms of pronunciation accuracy, tone correctness, and overall naturalness. This work represents a significant advancement in Hakka language technology and provides a valuable resource for language preservation and revitalization efforts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。