首个支持28种欧洲语言的开源语音-语言模型投影器,实现端到端多语种语音识别。
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

- 用Whisper编码器对接开源多语LLM,构建跨语言语音投影器
- 在高/低资源语言上均表现优异,28种语言端到端语音识别效果强
- 可快速扩展至新语言,且支持多语种翻译与话题识别
轻量级投影器是连接预训练语音编码器与大语言模型(LLM)的成熟方法,可将声学特征映射为词元级嵌入,用于语音识别(ASR)和语音问答等任务。然而现有系统通常仅支持少数语言,且多限于英语。本文提出MEUSLI,首个面向开放科学的多语言投影器家族,将Whisper编码器与开源多语LLM结合,实现28种欧洲语言的完全开源端到端语音识别。相比单语管道,MEUSLI在高资源和低资源语言上均表现强劲。通过适当的持续学习技术,可轻松扩展至训练中未见的新语言。此外,仅需每语言数小时的任务特定标注,即可将该投影器用于多语种语音翻译和话题识别。整体而言,MEUSLI为多语种语音理解任务提供了坚实基础,支持可扩展、包容性的开源语音大模型应用。
原文摘要 · Abstract (English)
Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering. Existing systems, however, typically only support a few languages and are often limited to English. We introduce MEUSLI, the first open-science multilingual projector family that links a Whisper encoder with open-source multilingual LLMs, enabling fully open-source end-to-end ASR in 28 European languages. MEUSLI extends prior monolingual pipelines, delivering strong results across high- and low-resource languages. Using proper continual leaning techniques, MEUSLI can be easily extended to other languages not seen in training. We further demonstrate that the MEUSLI projector can be leveraged beyond ASR, enabling multilingual speech translation and topic identification with only a few hours of task specific supervision per language. Overall, MEUSLI provides a solid foundation for multilingual speech understanding tasks, supporting scalable and inclu- sive open-source SpeechLLM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。