arXiv:2507.08477cs.CL2025-07中稿 · Interspeech 2025 M…被引 1

通过迭代式训练提升多语言语音识别效果,解决LoRA过拟合问题。

ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition

  • 分三阶段迭代优化:聚焦-反馈-修正,逐步提升模型性能。
  • 在国际挑战赛中获1、4名,验证方法有效性。
  • 适合需要高效微调的多语言语音识别场景。

大型语言模型与自动语音识别系统的深度融合已成为具有高实用价值的研究方向。为解决低秩适配(LoRA)在监督微调阶段常见的过拟合问题,本文提出一种创新的迭代式LoRA训练范式(ILT),结合迭代伪标签策略,有效提升了模型性能的理论上限。基于Whisper-large-v3和Qwen2-Audio,采用三阶段训练流程:聚焦训练、反馈训练和修正训练。实验结果表明该方法显著有效。此外,MegaAIS研究团队将该技术应用于Interspeech 2025多语言对话语音语言建模挑战赛(MLC-SLM),在Track 1(多语言语音识别任务)中获得第4名,在Track 2(语音分离与识别任务)中夺得第1名,充分展示了该方法的实践可行性与强应用潜力。

原文摘要 · Abstract (English)

The deep integration of large language models and automatic speech recognition systems has become a promising research direction with high practical value. To address the overfitting issue commonly observed in Low-Rank Adaptation (LoRA) during the supervised fine-tuning (SFT) stage, this work proposes an innovative training paradigm Iterative LoRA Training (ILT) in combination with an Iterative Pseudo Labeling strategy, effectively enhancing the theoretical upper bound of model performance. Based on Whisper-large-v3 and Qwen2-Audio, we conduct systematic experiments using a three-stage training process: Focus Training, Feed Back Training, and Fix Training. Experimental results demonstrate the effectiveness of the proposed method. Furthermore, the MegaAIS research team applied this technique in the Interspeech 2025 Multilingual Conversational Speech Language Modeling Challenge (MLC-SLM), achieving 4th in Track 1 (Multilingual ASR Task) and 1st place in Track 2 (Speech Separation and Recognition Task), showcasing the practical feasibility and strong application potential of our approach.

语音识别LoRA多语言微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。