用跨语言数据和噪声语音提升低资源语音合成质量
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
- 借由相关语言的高质量数据增强目标语言语音合成
- 用降噪模型处理非录音棚环境的低质ASR数据
- 通过知识蒸馏让小模型生成更清晰的语音,适合低资源语言
近年来,文本转语音(TTS)技术在英语上取得显著进展,主要得益于大规模、高质量的网络数据。然而,许多其他语言缺乏此类资源,仅依赖有限的录音棚级数据,导致合成语音常出现可懂度问题,尤其在低频字符二元组上表现更差。本文提出三种解决方案:首先,利用语义或地理相关的语言高质量数据来提升目标语言的TTS性能;其次,采用非录音棚环境下采集的低质量自动语音识别(ASR)数据,并通过去噪与语音增强模型进行清理;第三,通过大规模模型的知识蒸馏,结合合成数据生成更鲁棒的输出。在印地语上的实验表明,该方法显著降低了可懂度问题,经人工评估验证有效。本方法为缺乏高质量数据的语言提供了一种可行路径,实现资源共享与共同提升。
原文摘要 · Abstract (English)
Recent advancements in Text-to-Speech (TTS) technology have led to natural-sounding speech for English, primarily due to the availability of large-scale, high-quality web data. However, many other languages lack access to such resources, relying instead on limited studio-quality data. This scarcity results in synthesized speech that often suffers from intelligibility issues, particularly with low-frequency character bigrams. In this paper, we propose three solutions to address this challenge. First, we leverage high-quality data from linguistically or geographically related languages to improve TTS for the target language. Second, we utilize low-quality Automatic Speech Recognition (ASR) data recorded in non-studio environments, which is refined using denoising and speech enhancement models. Third, we apply knowledge distillation from large-scale models using synthetic data to generate more robust outputs. Our experiments with Hindi demonstrate significant reductions in intelligibility issues, as validated by human evaluators. We propose this methodology as a viable alternative for languages with limited access to high-quality data, enabling them to collectively benefit from shared resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。