arXiv:2606.13253cs.SDcs.AI2026-06

针对口吃语音识别,提出个性化联邦学习方法提升准确率。

Towards Personalized Federated Learning for Dysarthric Speech Recognition

  • 用参数与嵌入两种方式聚合模型,实现语音识别个性化
  • 在UASpeech和TORGO数据集上分别降低0.99%和0.56%错误率
  • 适合关注隐私保护下语音识别优化的研究者

口吃语音的语音识别极具挑战性。基于联邦学习(FL)的自动语音识别(ASR)虽能有效保护隐私,但因说话人差异导致的异构性问题影响性能。强制所有说话人共享相同模型组件在异构环境下表现不佳,因此个性化成为重要方向;然而针对口吃语音的相关研究仍较匮乏。本文探索两种聚合策略以实现个性化:基于参数的平均策略与基于嵌入的平均策略。在UASpeech和TORGO数据集上的实验表明,所提方法相比基线正则化FedAvg,在UASpeech上实现最高0.99%绝对(3.15%相对)的词错误率(WER)降低,在TORGO上实现0.56%绝对(4.73%相对)降低,统计显著。

原文摘要 · Abstract (English)

Speech recognition is challenging for dysarthric speakers. While federated learning (FL)-based ASR can be an effective tool for protecting privacy, it suffers from heterogeneity issues caused by speaker variability. Forcing all speakers to share the same model components can be suboptimal under such heterogeneity, making personalization a promising direction; however, related research on dysarthric speech remains limited. To this end, this paper explores two aggregation strategies to achieve personalization, including the parameter-based averaging strategy and the embedding-based averaging strategy. Experiments on UASpeech and TORGO show that the proposed methods outperform the baseline regularized FedAvg by statistically significant WER reductions of up to 0.99% absolute (3.15% relative) on UASpeech and 0.56% absolute (4.73% relative) on TORGO, respectively.

语音识别联邦学习个性化口吃语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。