直接语音到语音翻译模型综述,揭示技术进展与挑战
Direct Speech-to-Speech Neural Machine Translation: A Survey
- 不依赖中间文本,直接将语音转为目标语言语音
- 在真实场景中性能仍落后于级联模型,尤其在低资源语言上
- 适合语音翻译初学者及希望突破现有瓶颈的研究者
语音到语音翻译(S2ST)模型可将一种语言的语音内容无损转换为另一种语言的目标语音,对促进跨语言交流具有重要意义。近年来,研究者提出了直接式S2ST模型,该类方法无需经过中间文本生成阶段,具备更低的解码延迟,并能更好保留语调、情感等非语言特征。然而,直接S2ST在实现自然流畅的无缝沟通方面仍有不足,其性能在实际应用中仍显著低于级联式模型,尤其在真实场景和低资源语言上的表现有待提升。本文首次系统综述了直接S2ST模型的发展现状,涵盖数据集、应用场景、评估指标等问题,批判性分析了主流模型在基准数据集上的表现,并提出当前面临的挑战与未来研究方向。
原文摘要 · Abstract (English)
Speech-to-Speech Translation (S2ST) models transform speech from one language to another target language with the same linguistic information. S2ST is important for bridging the communication gap among communities and has diverse applications. In recent years, researchers have introduced direct S2ST models, which have the potential to translate speech without relying on intermediate text generation, have better decoding latency, and the ability to preserve paralinguistic and non-linguistic features. However, direct S2ST has yet to achieve quality performance for seamless communication and still lags behind the cascade models in terms of performance, especially in real-world translation. To the best of our knowledge, no comprehensive survey is available on the direct S2ST system, which beginners and advanced researchers can look upon for a quick survey. The present work provides a comprehensive review of direct S2ST models, data and application issues, and performance metrics. We critically analyze the models' performance over the benchmark datasets and provide research challenges and future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。