用正则化联邦学习提升失语和老年语音识别精度,兼顾隐私与性能。
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
- 在参数、嵌入和损失三层面设计正则化策略,缓解数据稀疏与异构问题。
- 在两个基准数据集上,相对词错误率降低最高达2.13%,绝对值降0.55%。
- 适合关注语音隐私保护的医疗与老龄化应用研究者。
目前准确识别失语及老年语音仍具挑战性。随着隐私保护需求推动从集中式方法转向联邦学习(FL)以保障数据安全,数据稀缺、分布不均与说话人差异等问题进一步加剧。为此,本文系统研究了用于隐私保护的失语与老年语音识别的正则化联邦学习技术,从参数、嵌入到新型损失三个层面进行优化。在基准数据集UASpeech(失语语音)与DementiaBank Pitt(老年语音)上的实验表明,正则化联邦学习系统相比基线FedAvg,在词错误率(WER)上实现统计显著的降低,绝对值最高达0.55%(相对降低2.13%)。进一步将通信频率提升至每批次一次,可接近集中式训练性能。
原文摘要 · Abstract (English)
Accurate recognition of dysarthric and elderly speech remains challenging to date. While privacy concerns have driven a shift from centralized approaches to federated learning (FL) to ensure data confidentiality, this further exacerbates the challenges of data scarcity, imbalanced data distribution and speaker heterogeneity. To this end, this paper conducts a systematic investigation of regularized FL techniques for privacy-preserving dysarthric and elderly speech recognition, addressing different levels of the FL process by 1) parameter-based, 2) embedding-based and 3) novel loss-based regularization. Experiments on the benchmark UASpeech dysarthric and DementiaBank Pitt elderly speech corpora suggest that regularized FL systems consistently outperform the baseline FedAvg system by statistically significant WER reductions of up to 0.55\% absolute (2.13\% relative). Further increasing communication frequency to one exchange per batch approaches centralized training performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。