用少数据实现低延迟长句语音翻译,关键在统一轨迹监督。
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

- 通过联合文本与声学语义码的轨迹监督,统一翻译路径。
- 仅需约2000小时数据,在90%数据缩减下仍保持稳定性能。
- 适合追求低延迟、少标注数据的实时语音翻译研发者。
长时序流式语音到语音翻译(S2ST)需在严格延迟约束下实现增量、无界翻译。现有方法常受限于句子级监督或依赖大量成对数据。本文提出一种训练方案,仅需约2000小时跨语言成对S2ST数据,结合辅助任务监督,即可实现句子级与长序列流式S2ST。基于多任务辅助训练,系统在配对数据减少90%时仍保持鲁棒性。核心创新为联合文本-声学语义码轨迹监督,将目标文本与声学语义码统一为一致承诺路径,无需独立不稳定的语音端发射控制器。此外,双流思维-说话者分解结构通过解耦语言推理与密集声学预测,有效缓解模态干扰,显著优于统一解码器基线。最终系统在RealSI和ACL60/60-dev上达到具有竞争力的质量-延迟权衡,其ASR-BLEU表现媲美闭源先进系统LiveInterpret~2.0。
原文摘要 · Abstract (English)
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only $\sim$2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90\%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。