arXiv:2507.17527cs.CLcs.SD2025-07被引 8

端到端实时语音翻译,克隆你的声音,延迟降至3秒。

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

  • 构建双向语音理解生成框架,实现高保真同步翻译。
  • 人工评测准确率超70%,翻译质量优于商用方案。
  • 克隆语音延迟从10秒降至3秒,适合会议、直播等场景。

同步口译(SI)是翻译行业最艰巨的挑战之一,现有自动系统长期面临转录与翻译质量差、无法实时生成语音、多说话人混淆及长篇发言中译文膨胀等问题。本文提出Seed-LiveInterpret 2.0,一个端到端的同步口译模型,支持高保真、超低延迟的语音到语音翻译,并具备语音克隆能力。作为全功能产品级解决方案,该模型通过创新的双工语音理解-生成框架,结合大规模预训练与强化学习,在复杂场景下实现翻译准确率与延迟的显著平衡,经人工译员评测,正确率超过70%。实验表明,其翻译质量显著优于商业方案,同时克隆语音平均延迟从近10秒降低至近实时的3秒,降幅约70%,极大提升实际可用性。

原文摘要 · Abstract (English)

Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.

语音翻译实时生成语音克隆端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。