arXiv:2603.16924eess.AScs.AI2026-03被引 1

无需训练,实时翻译长段语音,突破同步语音翻译瓶颈。

SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation

  • 利用预训练模型的交叉注意力机制管理输入历史与输出选择
  • 在8种语言的MuST-C数据集上达到媲美端到端模型的质量-延迟平衡
  • 适合需快速部署、无训练资源的实时多语种会议与直播场景

同步语音到语音翻译(SimulS2S)对实时多语言交流至关重要,正被越来越多地集成到会议和流媒体平台中。然而,该领域研究仍不充分,现有方法通常依赖资源密集型训练过程,且仅适用于短句预分段输入,难以泛化至连续语音。为此,我们提出首个无需训练的长段同步语音翻译策略 SimulU。SimulU 采用历史管理与语音输出选择策略,利用预训练端到端模型中的交叉注意力机制,调控输入历史与输出生成。在包含8种语言的 MuST-C 数据集上的评估表明,SimulU 在质量-延迟权衡上优于或相当于强基准级级联模型。通过消除额外训练需求,SimulU 为真实场景下的端到端长段同步语音翻译提供了可行路径。

原文摘要 · Abstract (English)

Simultaneous speech-to-speech translation (SimulS2S) is essential for real-time multilingual communication, with increasing integration into meeting and streaming platforms. Despite this, SimulS2S remains underexplored in research, where current solutions often rely on resource-intensive training procedures and operate on short-form, pre-segmented utterances, failing to generalize to continuous speech. To bridge this gap, we propose SimulU, the first training-free policy for long-form SimulS2S. SimulU adopts history management and speech output selection strategies that exploit cross-attention in pre-trained end-to-end models to regulate both input history and output generation. Evaluations on MuST-C across 8 languages show that SimulU achieves a better or comparable quality-latency trade-off against strong cascaded models. By eliminating the need for ad-hoc training, SimulU offers a promising path to end-to-end SimulS2S in realistic, long-form scenarios.

语音翻译同步翻译零训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。