arXiv:2606.13121cs.CLcs.AI2026-06

让实时语音翻译更自然,减少打断性停顿。

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

论文配图:NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation
图 1 · 摘自论文原文
  • 利用语言多样性和时长变化信号,优化翻译节奏
  • 显著减少段落间停顿,保持低延迟和高翻译质量
  • 适合需要流畅交互的实时翻译场景

同步语音到语音翻译旨在通过降低延迟实现近实时沟通,是连续翻译高延迟问题的有力替代方案。然而,过度追求低延迟常导致分段式语音输出,使听者面临频繁停顿的不自然听觉体验,增加认知负担。为此,我们提出一种流畅性感知优化框架,旨在在同步翻译的低延迟优势与连续翻译的自然语流之间找到平衡点。该框架通过利用模型内部信号——包括语言多样性与诱导的时间波动性——最小化段间静音。在短句和长句基准上的实验表明,该方法在保持竞争性延迟和翻译质量的同时,有效提升了语音流的自然性。

原文摘要 · Abstract (English)

Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation. However, the excessive pursuit of low latency often results in fragmented chunk-wise speech. Consequently, listeners are subjected to an unnatural acoustic flow punctuated by frequent pauses, which could increase their cognitive load. To bridge this gap, we introduce a fluency-aware optimization framework designed to discover the sweet spot between the low-latency benefits of simultaneous translation and the natural flow of consecutive translation. Our framework minimizes inter-chunk silences by leveraging model-internal signals, including linguistic diversity and induced temporal variability in speech durations. Experiments on short- and long-form benchmarks show that our framework produces natural speech flow while maintaining competitive latency and translation quality.

语音翻译实时交互流畅性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。