arXiv:2603.06310eess.AScs.CL2026-03中稿 · Interspeech 2026被引 2

针对太平洋原住民语言数据稀缺问题,提出持续适配新方法。

Continual Adaptation for Pacific Indigenous Speech Recognition

  • 采用低秩适应(LoRA)与持续学习框架应对语言差异
  • 三类语言实验显示模型内部表征严重漂移
  • 适合研究低资源语言语音识别的学者参考

语音基础模型在低资源太平洋原住民语言上表现不佳,主要因数据极度匮乏。此外,全量微调易引发灾难性遗忘。为此,本文开展实证研究,将模型适配至真实世界中的太平洋语种数据集。我们考察了数据量、适配策略及表征漂移对多种太平洋语言的影响,并分析了连续学习框架下顺序语言习得的效果。在三种不同太平洋原住民语言上的实验表明,适配语言距离较远的语言会引发严重的内部表征漂移,导致模型面临严格的可塑性与稳定性权衡。尽管LoRA在初期适配效果良好,但在连续学习过程中仍出现灾难性遗忘。该研究凸显了为未充分代表语言设计鲁棒适配策略的紧迫性。

原文摘要 · Abstract (English)

Speech foundation models struggle with low-resource Pacific Indigenous languages because of severe data scarcity. Furthermore, full fine-tuning risks catastrophic forgetting. To address this gap, we present an empirical study adapting models to real-world Pacific datasets. We investigate the impact of data volume, adaptation strategies, and representational drift on speech foundation models for various Pacific languages. Additionally, we analyze a continual learning framework for sequential language acquisition. Empirical results across three distinct Pacific Indigenous languages demonstrate that adapting to these linguistically distant languages induces severe internal representational drift. Consequently, these models face a strict plasticity and stability dilemma. While LoRA adapts well initially, it suffers from catastrophic forgetting during sequential learning. Ultimately, this study highlights the urgent need for robust adaptation strategies tailored to underrepresented languages.

语音识别低资源语言持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。