arXiv:2507.17288cs.CLeess.AS2025-07中稿 · Interspeech 2025 M…被引 1

用大模型提升多语言对话语音识别准确率

Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge

  • 设计编码器-适配器-大模型架构,融合语言模型推理能力
  • 在多语言对话数据上分阶段训练,显著降低错误率
  • 在挑战赛中获第二名,适合多语言语音系统研究者

本文介绍了我们提交至多语言对话语音建模挑战赛(MLC-SLM Challenge)任务1的Triple X语音识别系统。研究聚焦于通过创新的编码器-适配器-大语言模型(LLM)架构,优化多语言对话场景下的语音识别准确率。该框架利用文本型大语言模型的强大推理能力,并融入领域特定适应机制。为进一步提升多语言识别性能,我们采用精心设计的多阶段训练策略,基于大规模多语言音频数据集进行训练。实验结果表明,该方法在开发集和测试集上均达到具有竞争力的词错误率(WER),在挑战赛排名中获得第二名。

原文摘要 · Abstract (English)

This paper describes our Triple X speech recognition system submitted to Task 1 of the Multi-Lingual Conversational Speech Language Modeling (MLC-SLM) Challenge. Our work focuses on optimizing speech recognition accuracy in multilingual conversational scenarios through an innovative encoder-adapter-LLM architecture. This framework harnesses the powerful reasoning capabilities of text-based large language models while incorporating domain-specific adaptations. To further enhance multilingual recognition performance, we adopted a meticulously designed multi-stage training strategy leveraging extensive multilingual audio datasets. Experimental results demonstrate that our approach achieves competitive Word Error Rate (WER) performance on both dev and test sets, obtaining second place in the challenge ranking.

语音识别多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。