arXiv:2604.09916cs.LGeess.AS2026-04

改进语音翻译延迟与质量平衡,提升实时效率最多7.1%。

Regularized Entropy Information Adaptation with Temporal-Awareness Networks for Simultaneous Speech Translation

论文配图:Regularized Entropy Information Adaptation with Temporal-Awareness Networks for Simultaneous Speech Translation
图 1 · 摘自论文原文
  • 引入监督对齐与时间感知网络,增强读写策略的时序理解。
  • 在Whisper模型上实现最高7.1%的流式效率提升(NoSE)。
  • REINA-TAN更优效率,REINA-SAN更抗读取循环问题。

同时语音翻译(SimulST)需在翻译质量与低延迟间取得平衡。现有方法REINA基于信息增益训练读/写策略,但其缺乏时序上下文,常导致过度读取音频才开始翻译。本文提出两种改进:监督对齐网络(REINA-SAN)和时间步增强网络(REINA-TAN)。实验表明,二者均显著优于基线并解决稳定性问题;其中,REINA-TAN在流式效率上达到更优帕累托前沿,而REINA-SAN对‘读取循环’更具鲁棒性。应用于Whisper模型后,两者在归一化流式效率(NoSE)指标上相比现有竞争基线最高提升7.1%。

原文摘要 · Abstract (English)

Simultaneous Speech Translation (SimulST) requires balancing high translation quality with low latency. Recent work introduced REINA, a method that trains a Read/Write policy based on estimating the information gain of reading more audio. However, we find that information-based policies often lack temporal context, leading the policy to bias itself toward reading most of the audio before starting to write. We improve REINA using two distinct strategies: a supervised alignment network (REINA-SAN) and a timestep-augmented network (REINA-TAN). Our results demonstrate that while both methods significantly outperform the baseline and resolve stability issues, REINA-TAN provides a slightly superior Pareto frontier for streaming efficiency, whereas REINA-SAN offers more robustness against 'read loops'. Applied to Whisper, both methods improve the pareto frontier of streaming efficiency as measured by Normalized Streaming Efficiency (NoSE) scores up to 7.1% over existing competitive baselines.

语音翻译流式处理效率优化深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。