用模型融合提升失语语音识别准确率,效果优于传统微调。
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
- 通过合并多个微调模型提升语音识别泛化能力
- 长音频识别错误率降低16.2%,整体相对减少12%
- 无需额外推理开销,适合数据少或不同架构场景
自动语音识别(ASR)随着语音基础模型(SFMs)取得进展,但在失语语音上表现下降,主要因发音变异大且数据稀缺。本研究针对语音可及性挑战赛,以Whisper为基础模型,探索模型融合方法提升ASR泛化能力。对比单路径微调合并与多轮独立训练后合并的策略,最佳多轮合并方案在整体上实现相对12%的词错误率(WER)下降,在长音频上达16.2%的相对降低,显著缓解了长音频这一主要损失来源。随着合并模型数量增加,性能持续提升,且在低数据环境下依然有效,跨不同模型架构具有良好泛化性。结果表明,模型融合是一种可复现、无需额外推理成本或超参调优的稳定改进方法。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility challenge, explored model merging to improve ASR generalization using Whisper as the base SFM. We compared fine-tuning with single-trajectory merging, combining models from one fine-tuning path, and multi-run merging, merging independently trained models. Our best multi-run merging approach achieved a 12% relative decrease of WER over classic fine-tuning, and a 16.2% relative decrease on long-form audios, a major loss contributor in dysarthric ASR. Merging more and more models led to continuous gains, remained effective in low-data regimes, and generalized across model architectures. These results highlight model merging as an easily replicable adaptation method that consistently improves ASR without additional inference cost or hyperparameter tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。