解决多语言低资源语音识别中的遗忘问题,让模型更均衡地学习各种语言。
Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR

- 用语言平衡的参考梯度,在统一空间中约束参数更新。
- 在Whisper-large-v3上实现接近零的平均遗忘率。
- 适合需要持续学习多语言低资源语音模型的研究者。
大规模预训练语音识别模型(如Whisper)具备强大的多语言能力,但在低资源语言上微调时易发生灾难性遗忘。尽管持续学习可缓解此问题,现有方法在多语言场景下难以控制跨任务干扰,主导语言会扭曲优化方向。本文提出统一梯度投影(UGP),通过语言平衡回放生成的参考梯度,在统一投影空间中约束参数更新。该方法在投影空间中均衡各语言贡献,降低主导语言偏差,提升跨语言稳定性。进一步表明,梯度级投影与数据级回放结合可协同提升稳定性和适应性。在多种低资源语言组和不同模型规模下,UGP均能有效适应并显著缓解遗忘。在Whisper-large-v3上,实现近零平均遗忘。
原文摘要 · Abstract (English)
Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue, existing methods struggle to regulate cross-task interference in multilingual settings, where dominant languages bias optimization. We propose Unified Gradient Projection (UGP), which constrains parameter updates using reference gradients from language-balanced replay in a unified projection space. By equalizing per-language contributions in the projection, UGP reduces dominant-language bias and improves cross-lingual stability. We further show that combining gradient-level projection with data-level replay yields complementary gains in stability and plasticity. Across diverse low-resource language groups and model scales, UGP enables effective adaptation while substantially mitigating forgetting. On Whisper-large-v3, it achieves near-zero average forgetting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。