arXiv:2501.19010cs.CLcs.SD2025-01NAACL被引 7

通过动态音素级对比学习,提升口吃语音识别准确率

DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition

  • 按音素分段进行动态对比学习,捕捉语音细微差异
  • 在UASpeech数据集上降低22.10%词错误率
  • 适合处理严重程度各异的口吃语音识别任务

口吃语音识别常因口吃严重程度内在差异及与正常语音的外部差异导致性能下降。为此,我们提出动态音素级对比学习(DyPCL)方法,获得跨不同说话人的不变表示。将语音语句分解为音素片段,实现音素级对比学习,并利用动态连接时序分类对齐。不同于以往关注语句级嵌入的研究,本方法通过细粒度学习可区分语音中细微部分。此外,引入动态课程学习,根据音素发音相似性,逐步从易混淆负样本过渡到难分辨负样本。按难度层级训练有效缓解了说话人固有变异性,更精准识别困难语音。在UASpeech数据集上评估,相比基线模型,整体口吃组平均词错误率降低22.10%。

原文摘要 · Abstract (English)

Dysarthric speech recognition often suffers from performance degradation due to the intrinsic diversity of dysarthric severity and extrinsic disparity from normal speech. To bridge these gaps, we propose a Dynamic Phoneme-level Contrastive Learning (DyPCL) method, which leads to obtaining invariant representations across diverse speakers. We decompose the speech utterance into phoneme segments for phoneme-level contrastive learning, leveraging dynamic connectionist temporal classification alignment. Unlike prior studies focusing on utterance-level embeddings, our granular learning allows discrimination of subtle parts of speech. In addition, we introduce dynamic curriculum learning, which progressively transitions from easy negative samples to difficult-to-distinguishable negative samples based on phonetic similarity of phoneme. Our approach to training by difficulty levels alleviates the inherent variability of speakers, better identifying challenging speeches. Evaluated on the UASpeech dataset, DyPCL outperforms baseline models, achieving an average 22.10\% relative reduction in word error rate (WER) across the overall dysarthria group.

语音识别口吃语音对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。