通过正则化视角优化语音翻译多任务学习,提升数据稀缺下的性能。
Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task
- 从正则化角度设计多任务学习框架,整合跨模态与同模态正则化机制。
- 在MuST-C数据集上,调优超参数达到接近顶尖的翻译性能。
- 适合研究语音翻译、多任务学习及正则化方法的开发者参考。
端到端语音到文本翻译通常面临配对语音-文本数据稀缺的问题。一种解决方法是利用机器翻译任务的双语数据进行多任务学习(MTL)。本文从正则化视角重新审视MTL,探索序列在模态内与模态间如何被正则化。通过深入分析一致性正则化(跨模态)与R-drop(同模态)的作用,揭示了它们对总正则化的分别贡献。同时证明,机器翻译损失的系数在MTL设置中也构成另一种正则化来源。基于这三种正则化源,本文提出高维空间中的最优正则化轮廓——正则化边界。实验表明,在正则化边界内调优超参数,可在MuST-C数据集上实现接近最先进水平的性能。
原文摘要 · Abstract (English)
End-to-end speech-to-text translation typically suffers from the scarcity of paired speech-text data. One way to overcome this shortcoming is to utilize the bitext data from the Machine Translation (MT) task and perform Multi-Task Learning (MTL). In this paper, we formulate MTL from a regularization perspective and explore how sequences can be regularized within and across modalities. By thoroughly investigating the effect of consistency regularization (different modality) and R-drop (same modality), we show how they respectively contribute to the total regularization. We also demonstrate that the coefficient of MT loss serves as another source of regularization in the MTL setting. With these three sources of regularization, we introduce the optimal regularization contour in the high-dimensional space, called the regularization horizon. Experiments show that tuning the hyperparameters within the regularization horizon achieves near state-of-the-art performance on the MuST-C dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。