arXiv:2603.11578cs.CL2026-03被引 1

无需预设策略的端到端语音翻译模型,实时生成译文与转录。

Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

  • 用概率等待标记动态控制读写时机,实现无策略端到端流式处理。
  • 在英日、德、俄语种上低延迟与高延迟场景均刷新最佳BLEU分数。
  • 适合需要低延迟实时翻译的场景,如会议同传或直播字幕。

同步机器翻译(SiMT)传统上依赖离线翻译模型与人工设计规则或学习策略。本文提出 Hikari,一种无需策略、完全端到端的模型,通过将读取/写入决策编码为概率等待标记机制,实现语音到文本的同步翻译与流式转录。我们还引入解码时间膨胀机制,降低自回归开销并保证训练分布均衡。此外,提出监督微调策略,使模型具备延迟恢复能力,显著提升质量-延迟权衡。在英语到日语、德语和俄语的评测中,Hikari 在低延迟与高延迟场景下均取得新的最优 BLEU 分数,超越近期基线方法。

原文摘要 · Abstract (English)

Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hikari, a policy-free, fully end-to-end model that performs simultaneous speech-to-text translation and streaming transcription by encoding READ/WRITE decisions into a probabilistic WAIT token mechanism. We also introduce Decoder Time Dilation, a mechanism that reduces autoregressive overhead and ensures a balanced training distribution. Additionally, we present a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off. Evaluated on English-to-Japanese, German, and Russian, Hikari achieves new state-of-the-art BLEU scores in both low- and high-latency regimes, outperforming recent baselines.

语音翻译流式处理端到端多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。