用少样本生成可控普通话口吃语音,提升语音识别鲁棒性。
DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation

- 基于离散语音标记,通过显式口吃标签控制重复、延长等模式。
- 仅用不到50小时数据微调,合成语音质量与可控性均超现有方法。
- 生成的口吃语音有效提升多个ASR模型性能,尤其适配低资源场景。
口吃语音识别仍具挑战,重复、延长和阻塞等不流畅现象破坏语音连续性与声学特征。这一问题在汉语场景中尤为突出,因口吃语音数据稀缺,难以训练应对多样不流畅模式的鲁棒语音识别模型。为此,本文提出DisSpeech,一种基于离散语音标记的低资源可控普通话口吃语音合成与语音识别数据增强框架。该框架引入显式口吃事件标签以控制不同不流畅模式,将文本与口吃事件标签通过非自回归掩码生成变压器映射为语义语音标记,并结合显式音高与能量建模实现韵律感知的声学重建。仅需不足50小时的普通话口吃语音进行微调,DisSpeech即可生成高质量且可控的口吃语音。实验表明,该方法在语音质量和事件可控性方面均优于以往口吃语音合成方法。此外,合成的口吃语音显著提升多个语音识别模型性能,其中Qwen3-ASR-0.6B在评估的普通话口吃语音识别任务上达到4.19%的最优词错误率(CER),对流利语音识别影响极小。
原文摘要 · Abstract (English)
Stuttered speech recognition remains challenging, with disfluencies such as repetitions, prolongations, and blocks disrupting speech continuity and acoustic patterns. This problem is further aggravated in Mandarin scenarios by the limited availability of stuttered speech data, which makes it difficult to train robust ASR models for diverse disfluency patterns. To address this problem, this paper proposes DisSpeech, a discrete speech token-based framework for low-resource controllable Mandarin stuttered speech synthesis and ASR data augmentation. The proposed framework introduces explicit stuttering event labels to control different disfluency patterns. Text and stuttering event labels are mapped into semantic speech tokens by a non-autoregressive masked generative Transformer, followed by prosody-aware acoustic reconstruction with explicit pitch and energy modeling. With fine-tuning using less than 50 hours of Mandarin stuttered speech, DisSpeech can generate controllable stuttered speech with competitive speech quality. Experimental results show that the proposed method outperforms previous stuttered speech synthesis methods in both speech quality and event controllability. Furthermore, the synthesized stuttered speech effectively improves multiple ASR models, with Qwen3-ASR-0.6B achieving a state-of-the-art CER of 4.19% on the evaluated Mandarin stuttered speech recognition task, while causing only slight degradation on fluent speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。