arXiv:2601.18904cs.SDcs.AI2026-01

让语音大模型无需重新训练就能适应小语种,靠少量示范样例实时调整。

MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning

  • 用高资源语音数据后训练增强模型推理时的上下文学习能力
  • 在儿童语音识别、跨语言翻译等任务上显著提升低资源场景表现
  • 适合想低成本适配小语种或罕见口音的开发者使用

生成式语音与音频AI需服务全球多语言、多文化用户,但现有语音大模型仍主要基于高资源数据训练。在低资源场景下,目标语言或说话人缺乏足够标注数据,直接微调易受领域偏移影响。本文提出Meta Speech In-Context Learning(MetaSICL),一种仅使用丰富高资源语音数据的后训练方法,增强语音大模型在推理时利用少量本地示例进行自适应的能力。尽管未接触过目标低资源领域,MetaSICL在两个主干模型上均提升了儿童语音识别、音频理解/推理及未见语言方向的语音翻译与识别性能。当有少量领域内数据时,以MetaSICL作为预热再进行强化学习,优于直接微调,在五种类型多样的低资源语言语音识别中表现最优。该方法为构建可全球化的语音大模型提供了实用路径。

原文摘要 · Abstract (English)

Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data. Globalizing such systems requires handling low-resource settings, where the target speakers, languages, or tasks are poorly represented in training data. In these regimes, collecting enough labeled in-domain data is often impractical, and the small corpora available may still under-represent the test distribution, making direct fine-tuning brittle under domain shift. In-Context Learning (ICL) offers an alternative: instead of updating model parameters for every underserved community, an auditory LLM can adapt at inference time by conditioning on a few local demonstrations. However, vanilla speech ICL remains limited because most auditory LLMs are not explicitly trained to use such demonstrations effectively. We address this gap with Meta Speech In-Context Learning (MetaSICL), a post-training recipe that strengthens an auditory LLM's in-context adaptation ability using only abundant high-resource speech data. Although MetaSICL never trains on the target low-resource domains, it improves performance across two backbones on children's ASR, audio understanding/reasoning, and speech translation and ASR in directions and languages unseen in post-training. We further study the case where some in-domain data is available, using low-resource language ASR as a case study, since recognition for underserved languages is central to globalizing generative AI. Here, using MetaSICL as a warmup for in-domain reinforcement learning yields the strongest results, outperforming direct fine-tuning across five typologically diverse languages. Overall, MetaSICL offers a practical route toward globalizing auditory LLMs by building inference-time adaptation into the model.

语音大模型低资源语言上下文学习推理适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。