直接语音翻译让跨语言对话更自然流畅,减少延迟和错误传播。
Direct Speech to Speech Translation: A Review
- 跳过文本中间环节,直接从语音到语音翻译
- 保留说话人音色与语调,提升翻译自然度
- 适合实时多语言沟通场景,如外交与旅游
语音到语音翻译(S2ST)是一项变革性技术,可弥合全球沟通鸿沟,实现在外交、旅游和国际贸易中的实时多语言交互。本文回顾S2ST的发展历程,对比传统级联模型(依赖自动语音识别、机器翻译和文本转语音组件)与新兴的端到端及直接语音翻译(DST)模型。级联模型虽模块化且组件可优化,但存在错误传播、延迟高、韵律丢失等问题;而直接S2ST模型通过保留语音特征与韵律,维持说话人身份,降低延迟,提升翻译自然度。然而,其仍受限于数据稀疏、计算成本高及低资源语言泛化能力不足。本文批判性评估两类方法及其权衡,提出未来改进实时多语言通信的方向。
原文摘要 · Abstract (English)
Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines the evolution of S2ST, comparing traditional cascade models which rely on automatic speech recognition (ASR), machine translation (MT), and text to speech (TTS) components with newer end to end and direct speech translation (DST) models that bypass intermediate text representations. While cascade models offer modularity and optimized components, they suffer from error propagation, increased latency, and loss of prosody. In contrast, direct S2ST models retain speaker identity, reduce latency, and improve translation naturalness by preserving vocal characteristics and prosody. However, they remain limited by data sparsity, high computational costs, and generalization challenges for low-resource languages. The current work critically evaluates these approaches, their tradeoffs, and future directions for improving real time multilingual communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。