不用重训练,用任务算术扩展语音翻译语言对
Task Arithmetic for Language Expansion in Speech Translation
- 用任务算术合并多个单语言语音翻译模型
- 在两个数据集上提升最高达4.92的BLEU分数
- 可拓展到无训练数据的语言对,适合资源少场景
大语言模型进展推动了语音-文本多模态基础模型的发展,在指令微调的语音翻译(ST)任务中表现优异。但新增语言对需重新训练,成本高昂。为此,我们提出无需重训练的方法,通过任务算术从已有单对一语音翻译模型构建一对多系统。直接应用任务算术会导致目标语言混淆,因此引入语言控制模型以确保生成正确语言。在MuST-C和CoVoST-2上的实验显示,BLEU分数分别提升最高4.66和4.92,COMET得分提升8.87和11.83。此外,本框架还能通过任务类比,基于现有机器翻译(MT)和语音翻译(ST)模型合成缺乏成对数据或预训练模型的语言对,实现有效扩展。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has gained interest in speech-text multimodal foundation models, achieving strong performance on instruction-tuned speech translation (ST). However, expanding language pairs is costly due to re-training on combined new and previous datasets. To address this, we aim to build a one-to-many ST system from existing one-to-one ST systems using task arithmetic without re-training. Direct application of task arithmetic in ST leads to language confusion; therefore, we introduce an augmented task arithmetic method incorporating a language control model to ensure correct target language generation. Our experiments on MuST-C and CoVoST-2 show BLEU score improvements of up to 4.66 and 4.92, with COMET gains of 8.87 and 11.83. In addition, we demonstrate our framework can extend to language pairs lacking paired ST training data or pre-trained ST models by synthesizing ST models based on existing machine translation (MT) and ST models via task analogies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。