用多功能语音合成增强低资源语音识别效果
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
- 用高质量语音合成生成数据,弥补低资源语言数据不足
- 在多种低资源场景下均显著提升识别准确率
- 首次研究文本多样性对语音识别的促进作用
尽管自动语音识别(ASR)系统在大规模数据集上表现优异,但在方言、口音、少数族群语言及长尾关键词等低资源场景中仍表现不佳,这些领域具有重要实际意义。随着多功能、高自然度的文本转语音(TTS)模型的发展,利用其生成高质量语音进行数据增强,成为一种低成本且高效的提升ASR性能的方法。在前所未有的多样低资源数据集上进行的全面实验表明,该方法能持续带来显著性能提升,验证了通过多功能TTS模型增强低资源ASR的有效性与广泛适用性。此外,我们深入分析了合成语音数据的关键特性,包括文本多样性、说话人多样性及合成数据量,其中首次系统研究了文本多样性的作用。研究成果为基于TTS的数据增强在实际中的应用提供了有益指导,推动了低资源ASR的发展。
原文摘要 · Abstract (English)
While automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompassing dialects, accents, minority languages, and long-tail hotwords, domains with significant practical relevance. With the advent of versatile and powerful text-to-speech (TTS) models, capable of generating speech with human-level naturalness, expressiveness, and diverse speaker profiles, leveraging TTS for ASR data augmentation provides a cost-effective and practical approach to enhancing ASR performance. Comprehensive experiments on an unprecedentedly rich variety of low-resource datasets demonstrate consistent and substantial performance improvements, proving that the proposed method of enhancing low-resource ASR through a versatile TTS model is highly effective and has broad application prospects. Furthermore, we delve deeper into key characteristics of synthesized speech data that contribute to ASR improvement, examining factors such as text diversity, speaker diversity, and the volume of synthesized data, with text diversity being studied for the first time in this work. We hope our findings provide helpful guidance and reference for the practical application of TTS-based data augmentation and push the advancement of low-resource ASR one step further.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。