arXiv:2606.17820cs.CL2026-06

用双语微调提升低资源语音识别,通过语言标识符增强效果

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

论文配图:Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation
图 1 · 摘自论文原文
  • 输入文本前加语言标识符,让模型同时预测语言和转写
  • 语言识别准确时,双语微调显著提升低资源语音识别性能
  • 语言识别差时,推理阶段仍提供标识符可改善识别效果

本研究探讨双语微调对低资源语言自动语音识别(ASR)的影响。我们在九组语言上进行跨语言评估,涵盖多种语言家族和书写系统。训练时,在每个输入文本前添加语言识别标记,使模型学习区分两种语言;推理时,模型仅从语音输入中联合预测语言和转写。由于语言判断错误的文本表现出较低的ASR性能,我们还进行了后续实验,即在训练和推理阶段均提供语言识别标记。结果表明,当语言识别准确率较高时,双语微调具有明显优势;而在语言识别性能较差的情况下,推理阶段引入语言识别标记有助于提升ASR表现。

原文摘要 · Abstract (English)

This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a range of language families and writing systems. To distinguish the two languages, during training, we pre-pend each input text with a language identification token. At inference, the model jointly predicts both the language and transcription from the speech input alone. As texts for which the language is incorrectly determined show low ASR performance, we also conduct a follow-up experiment in which the language identification token is provided both during training and inference. Our results show that bilingual fine-tuning can be beneficial when language identification accuracy is high, and that in cases where language identification performance is low, including the language identification token at inference helps to improve ASR performance.

语音识别低资源双语微调语言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。