用合成数据提升印尼语语音识别对口吃语音的鲁棒性
Stuttering-Aware Automatic Speech Recognition for Indonesian Language
- 通过规则+大模型生成口吃语音,再用语音合成构建训练数据
- 在口吃语音上错误率下降,流畅语音性能不变
- 为低资源语言开发包容性语音技术提供新路径
自动语音识别系统在流畅语音上表现优异,但在口吃语音上性能显著下降,这一问题在印尼语等低资源语言中尤为突出,因缺乏专用数据集。为此,我们提出一种数据增强框架:结合规则变换与大语言模型,将重复和延音注入流畅文本,再通过文本到语音合成生成合成口吃音频。利用该合成数据微调预训练的印尼语Whisper模型,使模型在无需大规模真实录音的情况下适应非流利语音模式。实验表明,这种针对性的合成数据训练能持续降低口吃语音的识别错误率,同时保持流畅语音的性能,验证了合成数据流水线在发展代表性不足语言的包容性语音技术中的有效性。
原文摘要 · Abstract (English)
Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like Indonesian where specialized datasets are virtually non-existent. To overcome this scarcity, we propose a data augmentation framework that generates synthetic stuttered audio by injecting repetitions and prolongations into fluent text through a combination of rule-based transformations and large language models followed by text-to-speech synthesis. We apply this synthetic data to fine-tune a pre-trained Indonesian Whisper model using transfer learning, enabling the architecture to adapt to dysfluent acoustic patterns without requiring large-scale real-world recordings. Our experiments demonstrate that this targeted synthetic exposure consistently reduces recognition errors on stuttered speech while maintaining performance on fluent segments, validating the utility of synthetic data pipelines for developing more inclusive speech technologies in under-represented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。