arXiv:2505.20445cs.CLcs.AI2025-05被引 3

用上下文学习让大模型学会濒危语言的语音识别,效果媲美专用模型。

In-context Language Learning for Endangered Languages in Speech Recognition

  • 通过提供相关文本样例,让大模型在无监督下学习新语言
  • 概率驱动方法比传统指令方法提升语言建模与语音识别性能
  • 在四种濒危语言上达到或超越专用模型水平,适合低资源语言研究者

全球约有7000种语言,但当前大型语言模型仅支持其中一小部分。已有研究表明,大模型可通过上下文学习(ICL)在无监督数据情况下掌握新任务。本文将此探索延伸至语音识别领域,研究大模型是否能通过ICL学习未见过的低资源语言。在四种未被训练过的濒危语言上进行实验,结果表明:提供更多相关文本样本可显著提升语言建模与自动语音识别(ASR)性能。此外,基于概率的方法优于传统的指令驱动方法。最后发现,利用ICL的大模型在这些语言上的ASR表现可媲美甚至超越为特定语言专门训练的模型,同时保留原有语言模型能力。代码已公开。

原文摘要 · Abstract (English)

With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for certain tasks without supervised data. We extend this investigation to speech recognition, investigating whether LLMs can learn unseen, low-resource languages through in-context learning (ICL). With experiments on four diverse endangered languages that LLMs have not been trained on, we find that providing more relevant text samples enhances performance in both language modelling and Automatic Speech Recognition (ASR) tasks. Furthermore, we show that the probability-based approach outperforms the traditional instruction-based approach in language learning. Lastly, we show ICL enables LLMs to achieve ASR performance that is comparable to or even surpasses dedicated language models trained specifically for these languages, while preserving the original capabilities of the LLMs. Our code is publicly available.

语音识别低资源语言上下文学习濒危语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。