arXiv:2506.13596cs.CLcs.SD2025-06中稿 · Interspeech MLCSLM…被引 3

对比Qwen与Gemma在多语言语音大模型中的表现,提升识别准确率。

Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems

  • 用Whisper编码器+投影层+不同解码器构建多语言语音系统
  • 使用Gemma3-12B作解码器时测试平均字错率16.63%,优于Qwen2.5-7B的18.6%
  • 适合关注多语种语音识别与大模型融合的研究者

本文针对MLC-SLM Challenge 2025,提出一种结合微调Whisper-large-v3编码器、高效投影架构与多种解码器配置的多语言语音识别与语言建模系统。采用三阶段训练策略,逐步优化编码器、投影层与语言模型组件。在私有测试集上,以Gemma3-12B作为解码器时获得16.63%的平均字错率(WER)与字符错率(CER),而使用Qwen2.5-7B解码器时为18.6%。结果表明该系统在多语言场景下具备竞争力。

原文摘要 · Abstract (English)

This paper presents our system for the MLC-SLM Challenge 2025, focusing on multilingual speech recognition and language modeling with large language models (LLMs). Our approach combines a fine-tuned Whisper-large-v3 encoder with efficient projector architectures and various decoder configurations. We employ a three-stage training methodology that progressively optimizes the encoder, projector, and LLM components. Our system achieves competitive performance with a private test average WER/CER result of 16.63% using the Gemma3-12B and 18.6% using the Qwen2.5-7B as decoder-only language model.

语音识别多语言大模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。