arXiv:2505.22072cs.SDeess.AS2025-05中稿 · Interspeech 2025被引 3

动态路由专家模型实现失语语音识别零样本适配与实时处理

On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition

  • 通过实时预测路由参数动态组合适配专家
  • 相比基线模型字错误率降低最多6.36%(绝对1.34%)
  • 适合低资源场景下失语语音的快速适配

本文提出一种基于混合专家(MoE)的失语语音识别基础模型说话人适配框架。该方法支持零样本适配与实时处理,并融合领域知识。通过动态组合基于语音障碍严重程度和性别条件的适配专家,利用实时预测的说话人依赖路由参数实现高效适配。采用KL散度进一步增强专家间的多样性并提升对未见说话人的泛化能力。在UASpeech数据集上的实验表明,与未适配的HuBERT/WavLM基线相比,该方法在字错误率(WER)上实现了最高达1.34%绝对(6.36%相对)的显著降低;在不同说话人数据量条件下,相比批量模式适配,平均降低2.55%绝对(11.44%相对),推理速度提升最高达7倍。最终达到16.35%的最低已发表WER(极低可懂度情况下为46.77%)。

原文摘要 · Abstract (English)

This paper proposes a novel MoE-based speaker adaptation framework for foundation models based dysarthric speech recognition. This approach enables zero-shot adaptation and real-time processing while incorporating domain knowledge. Speech impairment severity and gender conditioned adapter experts are dynamically combined using on-the-fly predicted speaker-dependent routing parameters. KL-divergence is used to further enforce diversity among experts and their generalization to unseen speakers. Experimental results on the UASpeech corpus suggest that on-the-fly MoE-based adaptation produces statistically significant WER reductions of up to 1.34% absolute (6.36% relative) over the unadapted baseline HuBERT/WavLM models. Consistent WER reductions of up to 2.55% absolute (11.44% relative) and RTF speedups of up to 7 times are obtained over batch-mode adaptation across varying speaker-level data quantities. The lowest published WER of 16.35% (46.77% on very low intelligibility) is obtained.

失语语音MoE零样本适配实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。