用合成数据和模型正则化提升低资源语音翻译性能
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
- 通过自动生成翻译和语音增强合成数据,弥补低资源语言训练不足
- 北黎凡特阿拉伯语仅用合成数据的系统优于真实数据的级联系统
- 知识蒸馏与最小贝叶斯风险解码有效提升多任务表现,适合低资源场景
本文介绍基尔工业大学(KIT)在IWSLT2025低资源语音翻译赛道的参赛系统。针对班巴语、北黎凡特阿拉伯语和突尼斯阿拉伯语到英语的三组语言对,我们构建了级联式(ASR+MT)与端到端(E2E)语音翻译系统。基于预训练模型,采用多种微调策略高效利用有限资源。研究进一步探索合成数据与模型正则化的系统增强方法:通过机器翻译模型从语音识别数据生成翻译以扩充训练集;对缺乏平行语音翻译数据的北黎凡特阿拉伯语,仅使用合成数据训练的系统表现略优于基于真实数据的级联系统。同时,利用文本转语音模型从翻译数据生成合成语音,显著提升班巴语的语音识别与语音翻译性能。此外,引入内部知识蒸馏(intra-distillation),在多个预训练模型上持续提升各任务表现。最后,通过最小贝叶斯风险(Minimum Bayes Risk)解码融合级联与端到端系统,整体性能提升约1.5 BLEU点。
原文摘要 · Abstract (English)
This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST) systems for three language pairs: Bemba, North Levantine Arabic, and Tunisian Arabic into English. Building upon pre-trained models, we fine-tune our systems with different strategies to utilize resources efficiently. This study further explores system enhancement with synthetic data and model regularization. Specifically, we investigate MT-augmented ST by generating translations from ASR data using MT models. For North Levantine, which lacks parallel ST training data, a system trained solely on synthetic data slightly surpasses the cascaded system trained on real data. We also explore augmentation using text-to-speech models by generating synthetic speech from MT data, demonstrating the benefits of synthetic data in improving both ASR and ST performance for Bemba. Additionally, we apply intra-distillation to enhance model performance. Our experiments show that this approach consistently improves results across ASR, MT, and ST tasks, as well as across different pre-trained models. Finally, we apply Minimum Bayes Risk decoding to combine the cascaded and end-to-end systems, achieving an improvement of approximately 1.5 BLEU points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。