arXiv:2602.01008eess.AScs.CL2026-02ACL被引 7

针对低资源语言语音识别,提出按层分配适应能力的新方法

Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages

  • 根据各层语言特异性差异,动态分配模型适配资源
  • 在18种低资源语言上,参数减少80%且错误率降低29%
  • 适合资源受限场景下的多语言语音识别系统优化

近期语音基础模型在高资源语言的多语言自动语音识别(ASR)中表现优异,但在低资源语言上的适配仍面临数据稀缺与效率限制。全模型微调计算成本高且易过拟合,而参数高效方法如LoRA对各层均匀适配,忽视内部表征差异,影响效果与效率。我们分析多语言ASR模型,发现适配性呈U型分布:早期和晚期层具语言特异性需更多适配,中间层保留共享语义需较少适配。基于此,提出深度感知适配框架DAMA,按层角色分配适配能力。DAMA引入基于奇异值分解(SVD)的初始化以约束适配并保持U型模式,并采用冻结中间层基函数进一步提升效率。在两个基准数据集上的18种低资源语言上评估,DAMA在参数量减少80%的前提下达到或超越现有最佳准确率,极端数据稀缺下错误率降低29%,显著提升内存、训练时间与计算效率。

原文摘要 · Abstract (English)

Recent speech foundation models excel at multilingual automatic speech recognition (ASR) for high-resource languages, but adapting them to low-resource languages remains challenging due to data scarcity and efficiency constraints. Full-model fine-tuning is computationally expensive and prone to overfitting, while parameter-efficient methods like LoRA apply adaptation uniformly across layers, overlooking internal representations thus compromising effectiveness and efficiency. We analyze multilingual ASR models and reveal a U-shaped adaptability pattern: early and late layers are language-specific and require more adaptation, while intermediate layers retain shared semantics and need less. Building on this observation, we propose DAMA, a Depth-Aware Model Adaptation framework that allocates adaptation capacity according to each layer's role. DAMA also introduces Singular Value Decomposition (SVD)-based initialization to constrain adaptation and preserve the U-shaped pattern, as well as a frozen middle-layer basis for further efficiency. Evaluated on 18 low-resource languages across two benchmark datasets, DAMA matches or surpasses state-of-the-art accuracy with 80% fewer trainable parameters, achieves a 29% error reduction under extreme data scarcity, and significantly improves memory, training time, and computational efficiency over baselines. These results highlight the benefits of structure-aware adaptation for efficient, scalable multilingual ASR.

多语言语音识别低资源学习参数高效结构感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。