arXiv:2606.03948cs.CL2026-06

10亿参数模型实现多语种实时语音翻译,性能超同类方案。

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

  • 采用对齐注意力策略,在离线模型中实现同步翻译
  • 低延迟与高精度兼得,10亿参数模型表现领先
  • 支持50种语言互译,适合资源受限场景

我们基于最先进的对齐注意力策略,将离线直接语音到文本翻译模型Canary实现为同步翻译系统,并提交至IWSLT 2026同步语音翻译共享任务,涵盖捷克语→英语、英语→德语和意大利语。系统优势包括:(1)翻译质量高,在计算无感知仿真中,无论低延迟还是高延迟场景均优于同等规模基线;(2)计算开销小,模型仅含10亿参数;(3)多语言支持,覆盖25种源语言和25种目标语言。

原文摘要 · Abstract (English)

We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Shared task for Czech to English and English to German and Italian. The strengths of our system are: (1) high translation quality, outperforming similarly sized baselines both in low- and high-latency regimes in computationally unaware simulations; (2) low computational requirements, as the model has only 1B parameters; (3) multilinguality -- support of 25 source and 25 target languages.

语音翻译多语言轻量化同步翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。