arXiv:2601.07038cs.CL2026-01被引 1

用语言间任务向量融合,提升低资源语音识别效果

Task Arithmetic with Support Languages for Low-Resource ASR

  • 通过微调Whisper模型生成语言任务向量,线性组合高-低资源语言向量
  • 在23种低资源语言上实现最高10%的词错误率降低
  • 适合需要快速适配新语言的低资源语音识别场景

由于自动语音识别(ASR)在大量低资源语言中具有广泛应用前景,开发资源受限的方法至关重要。现有方法常利用与目标语言相近的高资源语言数据。本文将特定语言的训练视为一个任务,通过微调Whisper ASR系统生成任务向量。针对每对高-低资源语言,采用线性组合方式合并任务向量,并在低资源语言的验证集上优化下游词错误率。在23种低资源目标语言上的实验表明,该方法相比基线模型实现了最高10%的词错误率改善。

原文摘要 · Abstract (English)

The development of resource-constrained approaches to automatic speech recognition (ASR) is of great interest due to its broad applicability to many low-resource languages for which there is scant usable data. Existing approaches to many low-resource natural language processing tasks leverage additional data from higher-resource languages that are closely related to a target low-resource language. One increasingly popular approach uses task arithmetic to combine models trained on different tasks to create a model for a task where there is little to no training data. In this paper, we consider training on a particular language to be a task, and we generate task vectors by fine-tuning variants of the Whisper ASR system. For pairs of high- and low-resource languages, we merge task vectors via a linear combination which is optimized on the downstream word error rate on the low-resource target language's validation set. Across 23 low-resource target languages for which we evaluate this technique, we find consistent word error rate improvements of up to 10% compared to a baseline without our approach.

语音识别低资源任务向量Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。