arXiv:2505.21781cs.CL2025-05被引 4

GMU改进了低资源语音翻译模型,提升多语言翻译效果。

GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task

  • 用SeamlessM4T-v2微调端到端语音翻译,支持多语言
  • 直接微调在多数语言上表现最佳,初始化编码器提升新语言性能
  • 多任务训练略有帮助,适合低资源场景研究者

本文介绍GMU团队参与IWSLT 2025低资源语音翻译共享任务的系统。我们为除黎凡特阿拉伯语外的所有语言对训练了系统,基于SeamlessM4T-v2微调自动语音识别(ASR)、机器翻译(MT)和端到端语音翻译(E2E ST)模型。其中ASR与MT模型也用于构建级联式语音翻译系统。此外,我们探索了多种端到端训练范式,包括直接微调、多任务训练以及使用微调过的ASR和/或MT模型进行参数初始化。实验结果表明:(1)直接端到端微调能取得良好效果;(2)使用已微调的ASR编码器初始化可提升未在SeamlessM4T-v2上训练过的语言的翻译性能;(3)多任务训练有轻微增益。

原文摘要 · Abstract (English)

This paper describes the GMU systems for the IWSLT 2025 low-resource speech translation shared task. We trained systems for all language pairs, except for Levantine Arabic. We fine-tuned SeamlessM4T-v2 for automatic speech recognition (ASR), machine translation (MT), and end-to-end speech translation (E2E ST). The ASR and MT models are also used to form cascaded ST systems. Additionally, we explored various training paradigms for E2E ST fine-tuning, including direct E2E fine-tuning, multi-task training, and parameter initialization using components from fine-tuned ASR and/or MT models. Our results show that (1) direct E2E fine-tuning yields strong results; (2) initializing with a fine-tuned ASR encoder improves ST performance on languages SeamlessM4T-v2 has not been trained on; (3) multi-task training can be slightly helpful.

语音翻译低资源SeamlessM4T端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。