提出分阶段对齐训练法,提升多语言语音-文本模型的性能
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
- 分阶段训练:先对齐同语言表示,再动态激活大模型进行跨语言对齐
- 在多个数据集上超越传统方法,跨语言通用性与语言特异性平衡更好
- 适合做多语言语音识别、翻译系统的研究者和开发者参考
大型语言模型(LLMs)已从文本扩展至语音,催生支持语音识别、翻译与合成的语音大模型(SLMs)。核心挑战在于语音与文本表示的对齐,尤其在多语言场景下更为复杂。现有方法通常冻结LLM参数,仅训练编码器处理多语言数据,这迫使跨语言收敛,限制性能提升。本文提出渐进式对齐表示训练(PART),一种多阶段多任务框架,将同语言与跨语言对齐分离。跨语言训练阶段动态激活LLM参数,并后期引入文本任务以增强多语言理解能力。在CommonVoice 15、Fleurs、Wenetspeech和CoVoST2上的实验表明,PART显著优于传统方法;分析证实其能有效平衡语言特异性与跨语言泛化能力。结果验证了PART在多语言语音模态对齐中的有效性与普适性。
原文摘要 · Abstract (English)
Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is aligning speech and text representations, which becomes harder in multilingual settings. Existing methods often freeze LLM parameters and train encoders on multilingual data, but this forces cross-language convergence and limits performance. We introduce Progressive Alignment Representation Training (PART), a multi-stage and multi-task framework that separates within-language from cross-language alignment. During cross-language training, LLM parameters are dynamically activated, and text-based tasks are later introduced to enhance multilingual understanding. Experiments on CommonVoice 15, Fleurs, Wenetspeech, and CoVoST2 show that PART surpasses conventional approaches, with analysis confirming its ability to balance language-specific distinctions and cross-language generalization. These results demonstrate PART's effectiveness and generality for multilingual speech modality alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。