让自监督语音模型持续学新语言不遗忘,仅用2%参数量
Lamer-SSL: Layer-aware Mixture of LoRA Experts for Continual Multilingual Expansion of Self-supervised Models without Forgetting
- 分层分配LoRA专家,深层学语义,浅层保通用
- 结合回放机制,用极少量数据防遗忘
- 适合需要持续扩展多语言的语音系统
尽管表现优异,自监督语音模型在扩展新语言时往往泛化能力不足,且在持续训练中容易遗忘已学知识。为此,我们提出Lamer-SSL,一种参数高效的框架,融合分层专家混合(Lamer)模块与回放策略。Lamer模块可灵活平衡共享与语言特异性表征,分层专家分配将更多专家置于深层,以利用更丰富的语义信息。同时,回放策略仅用少量数据即可保留先前知识,缓解持续训练中的遗忘问题。在自动语音识别(ASR)和语言识别(LID)任务上的实验表明,Lamer-SSL能有效将自监督模型扩展至新语言,同时在旧语言上保持强性能,仅有2.14%的参数可训练。
原文摘要 · Abstract (English)
Despite their impressive performance, self-supervised speech models often struggle to generalize to new languages and tend to forget previously acquired knowledge during continual training. To address this, we propose Lamer-SSL, a parameter-efficient framework that integrates a Layer-Aware MixturE of LoRA Experts (Lamer) module with a replay strategy. The Lamer module enables flexible balancing between shared and language-specific representations, while layer-aware expert allocation assigns more experts to deeper layers where semantic information is richer. Meanwhile, the replay strategy retains prior knowledge using minimal data, mitigating forgetting during continual training. Experiments on automatic speech recognition (ASR) and language identification (LID) demonstrate that Lamer-SSL extends self-supervised models to new languages effectively while maintaining strong performance on previously learned languages with only 2.14% parameters being trainable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。