arXiv:2503.11080cs.CLcs.SD2025-03被引 1

提出多语言同步语音翻译联合训练与解码方法

Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation

  • 设计统一解码器架构支持多语言同步训练
  • 异步训练策略提升跨语言知识迁移效果
  • 构建多向对齐数据集用于真实场景评估

近期端到端语音翻译研究推动了多语言及同步语音翻译的发展。本文在更贴近实际应用的多对一多语言场景下,研究端到端同步语音翻译。探索了独立解码器与统一解码器架构,并在统一架构上提出异步训练策略以促进跨语言知识迁移。为此构建了一个多向对齐的多语言端到端语音翻译数据集作为基准测试平台。实验结果表明,所提模型在该数据集上表现有效。代码与数据已公开于:https://github.com/XiaoMi/TED-MMST。

原文摘要 · Abstract (English)

Recent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real scenarios. We explore a separate decoder architecture and a unified architecture for joint synchronous training in this scenario. To further explore knowledge transfer across languages, we propose an asynchronous training strategy on the proposed unified decoder architecture. A multi-way aligned multilingual end-to-end ST dataset was curated as a benchmark testbed to evaluate our methods. Experimental results demonstrate the effectiveness of our models on the collected dataset. Our codes and data are available at: https://github.com/XiaoMi/TED-MMST.

语音翻译多语言同步翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。