用LoRA专家提升多语言语音识别微调效率与精度。
Efficient Multilingual ASR Finetuning via LoRA Language Experts
- 为每种语言设计LoRA专家,通过融合或蒸馏实现高效微调。
- 在有语言意识和无语言意识场景下分别提升10%和15%性能。
- 适合需要定制化多语言语音识别的部署场景。
深度学习的进展推动了多语言自动语音识别(ASR)的发展,得益于先进的模型架构和大规模多语言数据集。然而,多语言模型仍面临多语言诅咒问题:不同语言相互干扰,难以有效识别多种语言,同时共享模型容量。本文提出一种基于Whisper的高效多语言ASR微调框架,通过预设的LoRA语言专家,结合专家融合或知识蒸馏方法,显著优于标准微调。实验表明,所提模型在语言感知和语言无关场景中分别实现约10%和15%的相对性能提升。
原文摘要 · Abstract (English)
Recent advancements in deep learning have significantly enhanced multilingual automatic speech recognition (ASR) due to the development of advanced model architectures and available large-scale multilingual datasets. Despite that, multilingual ASR still suffers from the curse of multilinguality in that different languages tend to interfere with each other, making it difficult for the ASR model to identify multiple languages effectively while sharing model capacity across them. This paper proposes an efficient finetuning framework for customized multilingual ASR via prepared LoRA language experts based on Whisper. Through LoRA expert fusion or knowledge distillation, our approach achieves better recognition performance on target languages than standard fine-tuning methods. Experimental results demonstrate that the proposed models yield approximately 10\% and 15\% relative performance gains in language-aware and language-agnostic scenarios, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。