直接语音到语音翻译模型Translatotron突破传统多阶段瓶颈
Speech to Speech Translation with Translatotron: A State of the Art Review
- 端到端直接转换语音,跳过语音识别和文本翻译步骤
- Translatotron3在部分指标上超越传统级联模型
- 特别适合非洲语言与主流语言间的翻译场景
语音到语音翻译长期依赖级联式方法(语音识别+文本翻译+语音合成),存在延迟高、错误累积等问题。Google提出的Translatotron系列模型通过端到端序列到序列直接翻译,解决上述缺陷。该系列包含三个版本:Translatotron 1作为概念验证,虽效果不如级联模型但表现可期;Translatotron 2显著改进,性能接近级联模型;最新版Translatotron 3在某些方面已超越级联模型。本文全面回顾语音到语音翻译技术,重点分析各版本Translatotron的演进与优势,并指出其在连接非洲语言与其他主流语言方面具有显著潜力。
原文摘要 · Abstract (English)
A cascade-based speech-to-speech translation has been considered a benchmark for a very long time, but it is plagued by many issues, like the time taken to translate a speech from one language to another and compound errors. These issues are because a cascade-based method uses a combination of methods such as speech recognition, speech-to-text translation, and finally, text-to-speech translation. Translatotron, a sequence-to-sequence direct speech-to-speech translation model was designed by Google to address the issues of compound errors associated with cascade model. Today there are 3 versions of the Translatotron model: Translatotron 1, Translatotron 2, and Translatotron3. The first version was designed as a proof of concept to show that a direct speech-to-speech translation was possible, it was found to be less effective than the cascade model but was producing promising results. Translatotron2 was an improved version of Translatotron 1 with results similar to the cascade model. Translatotron 3 the latest version of the model is better than the cascade model at some points. In this paper, a complete review of speech-to-speech translation will be presented, with a particular focus on all the versions of Translatotron models. We will also show that Translatotron is the best model to bridge the language gap between African Languages and other well-formalized languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。