针对濒危原住民语言数据少难题,提出两种新方法提升模型持续学习能力。
Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification

- 结合重放与权重保护,防止模型遗忘旧语言知识
- 在3种原住民语言上实现更高识别准确率
- 适合资源稀缺语言的长期模型迭代场景
语言识别是将濒危澳大利亚原住民语言(AALs)融入语音技术、支持语言复兴与数字包容的关键步骤。然而,极端数据稀缺严重制约模型性能。从高资源语言迁移学习虽有潜力,但适应新语言时常引发灾难性遗忘。持续学习(CL)可缓解此问题,但在极低数据条件下仍具挑战。为此,我们提出两种混合持续学习方法:增强型回放弹性权重巩固(Replay Augmented Elastic Weight Consolidation)与约束引导知识蒸馏(Constraint Guided Knowledge Distillation),以适配预训练语音模型进行AAL识别,同时保留先前学习的知识。在Warlpiri、Dalabon和Dharawal三个语种上的实验表明,所提方法优于微调及现有CL基线,在多语言适应中表现更优,且保持对已学高资源语言的性能。
原文摘要 · Abstract (English)
Language identification is an important step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies supporting language revitalisation and digital inclusion. However, extreme data scarcity limits model performance. Transfer learning from high-resource languages shows promise but often suffers from catastrophic forgetting when adapting to new languages. Continual learning (CL) can mitigate this issue, though it remains challenging with very limited data. To address this, we propose two hybrid continual learning methods: Replay Augmented Elastic Weight Consolidation and Constraint Guided Knowledge Distillation to adapt pretrained speech models for AAL identification while preserving previously learned knowledge. Experiments on Warlpiri, Dalabon and Dharawal show that the proposed methods outperform fine-tuning and existing CL baselines, improving adaptation to multiple AALs while maintaining performance on previously learnt high-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。