通过低秩稀疏合并,高效整合多语言语音模型并提升性能。
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
- 采用低秩与稀疏剪枝融合策略,保留关键参数结构。
- 跨10种语言实验显示性能提升超20%。
- 适合需要快速扩展多语言能力的语音系统开发者。
语言多样性给语音转文本(S2T)任务带来显著挑战,如自动语音识别和翻译。传统多语言多任务训练方法虽能联合优化多种语言的语音识别与翻译任务,但存在计算成本高、语言干扰严重、训练配置不佳及可扩展性差等问题。为此,我们提出LoRS-Merging(低秩与稀疏模型合并)技术,通过结合低秩表示与稀疏剪枝,在保留关键结构的同时去除冗余参数,有效缓解语言干扰并提升可扩展性。在10种语言上的实验表明,该方法显著优于多语言多任务训练、顺序训练及其他合并方法,性能提升超过20%。结果表明,模型合并,特别是LoRS-Merging,是传统多语言训练的有效补充,适用于可扩展的S2T应用。
原文摘要 · Abstract (English)
Language diversity presents a significant challenge in speech-to-text (S2T) tasks, such as automatic speech recognition and translation. Traditional multi-lingual multi-task training approaches aim to address this by jointly optimising multiple speech recognition and translation tasks across various languages. While models like Whisper, built on these strategies, demonstrate strong performance, they still face issues of high computational cost, language interference, suboptimal training configurations, and limited extensibility. To overcome these challenges, we introduce LoRS-Merging (low-rank and sparse model merging), a novel technique designed to efficiently integrate models trained on different languages or tasks while preserving performance and reducing computational overhead. LoRS-Merging combines low-rank and sparse pruning to retain essential structures while eliminating redundant parameters, mitigating language interference, and enhancing extensibility. Experimental results across 10 languages demonstrate that LoRS-Merging significantly outperforms multi-lingual multi-task training, sequential training, and other merging methods, achieving over 20% improvement in normalised performance. Our findings suggest that model merging, particularly LoRS-Merging, is a scalable and effective complement to traditional multi-lingual training strategies for S2T applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。