arXiv:2508.05149eess.AScs.AI2025-08中稿 · Interspeech 2025被引 11

用轻量投影器连接语音编码器与大模型,提升低资源语言识别性能。

Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages

  • 通过可训练的轻量投影器连接语音编码器与大语言模型,构建SLAM-ASR框架。
  • 在小样本数据下,预训练于高资源语言的多语言投影器显著缓解数据稀缺问题。
  • 适用于低资源语言语音识别研究,尤其适合关注跨语言迁移的研究者。

大型语言模型(LLMs)在高资源语言的语音输入处理中展现出潜力,达到多项任务的顶尖性能。然而其在低资源场景中的应用仍待探索。本文基于SLAM-ASR框架研究了面向低资源自动语音识别的语音大模型应用,该框架通过一个可训练的轻量级投影器连接语音编码器与大语言模型。首先,评估了达到Whisper-only性能所需的训练数据量,重申了数据匮乏的挑战。其次,表明利用在高资源语言上预训练的单语或多语投影器可显著降低数据稀缺的影响,尤其在小训练集情况下表现更优。使用多语言LLM(EuroLLM、Salamandra)与whisper-large-v3-turbo,在多个公开基准上进行评估,为未来优化低资源语言和多语言语音大模型提供了洞见。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource settings. This work investigates the use of Speech LLMs for low-resource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM. Firstly, we assess training data volume requirements to match Whisper-only performance, re-emphasizing the challenges of limited data. Secondly, we show that leveraging mono- or multilingual projectors pretrained on high-resource languages reduces the impact of data scarcity, especially with small training sets. Using multilingual LLMs (EuroLLM, Salamandra) with whisper-large-v3-turbo, we evaluate performance on several public benchmarks, providing insights for future research on optimizing Speech LLMs for low-resource languages and multilinguality.

语音识别低资源大模型迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。