用新方法生成更像口吃语音的合成数据,提升语音识别准确率。
DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement
- 基于匹配音高流模型,分阶段预测异常语调节奏。
- 合成语音与真实口吃语音相似度达MCD 4.29,误差降低54.22%。
- 适合做口吃语音识别增强的研究者或医疗语音系统开发者。
口吃语音存在异常语调和显著说话人差异,长期困扰自动语音识别(ASR)。尽管文本转语音(TTS)数据增强有潜力,但现有方法难以准确建模口吃语音的病理节奏与声学特征。为此,我们提出基于Matcha-TTS架构的DARS框架,包含多阶段节奏预测器(通过正常与口吃语音对比偏好优化)及口吃风格条件流匹配机制,协同提升时间节奏重建与病理声学风格模拟。在TORGO数据集上的实验表明,DARS实现均值倒谱距离(MCD)为4.29,接近真实口吃语音。将基于Whisper的ASR系统用DARS生成的合成口吃语音进行适配,相较当前最优方法实现54.22%的词错误率(WER)相对下降,验证了该框架在提升识别性能方面的有效性。
原文摘要 · Abstract (English)
Dysarthric speech exhibits abnormal prosody and significant speaker variability, presenting persistent challenges for automatic speech recognition (ASR). While text-to-speech (TTS)-based data augmentation has shown potential, existing methods often fail to accurately model the pathological rhythm and acoustic style of dysarthric speech. To address this, we propose DARS, a dysarthria-aware rhythm-style synthesis framework based on the Matcha-TTS architecture. DARS incorporates a multi-stage rhythm predictor optimized by contrastive preferences between normal and dysarthric speech, along with a dysarthric-style conditional flow matching mechanism, jointly enhancing temporal rhythm reconstruction and pathological acoustic style simulation. Experiments on the TORGO dataset demonstrate that DARS achieves a Mean Cepstral Distortion (MCD) of 4.29, closely approximating real dysarthric speech. Adapting a Whisper-based ASR system with synthetic dysarthric speech from DARS achieves a 54.22% relative reduction in word error rate (WER) compared to state-of-the-art methods, demonstrating the framework's effectiveness in enhancing recognition performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。