arXiv:2603.05354cs.CLeess.AS2026-03

用模型合并提升语音识别多领域适应能力,避免重复训练。

Exploring the potential and limitations of Model Merging for Multi-Domain Adaptation in ASR

  • 提出新算法BoostedTSV-M,缓解秩崩溃问题。
  • 在10个葡萄牙语领域上表现优于全量微调。
  • 适合需要多领域泛化能力的语音系统部署者。

模型合并是一种可扩展的多任务训练替代方案,能将多个专用模型的能力融合为单一模型。这对大型语音基础模型尤为关键,因其通常通过领域特定微调得到多个定制检查点,而新数据出现时重复全量微调计算成本过高。本文研究模型合并在多领域语音识别中的应用,针对10个欧洲葡萄牙语领域基准测试了11种合并算法,评估了域内准确率、分布偏移下的鲁棒性,以及英文和多语言性能。我们进一步提出基于TSV-M的BoostedTSV-M算法,通过奇异值增强缓解秩崩溃并提升数值稳定性。整体上,该方法在欧洲葡萄牙语上优于全量微调,同时保持跨域泛化能力。

原文摘要 · Abstract (English)

Model merging is a scalable alternative to multi-task training that combines the capabilities of multiple specialised models into a single model. This is particularly attractive for large speech foundation models, which are typically adapted through domain-specific fine-tuning, resulting in multiple customised checkpoints, for which repeating full fine-tuning when new data becomes available is computationally prohibitive. In this work, we study model merging for multi-domain ASR and benchmark 11 merging algorithms for 10 European Portuguese domains, evaluating in-domain accuracy, robustness under distribution shift, as well as English and multilingual performance. We further propose BoostedTSV-M, a new merging algorithm based on TSV-M that mitigates rank collapse via singular-value boosting and improves numerical stability. Overall, our approach outperforms full fine-tuning on European Portuguese while preserving out-of-distribution generalisation in a single model.

语音识别模型合并多领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。