用信息增益指导翻译延迟,提升实时语音翻译质量
REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation
- 基于信息论设计自适应等待策略,动态决定何时输出
- 在法语、西班牙语、德语间实现顶尖流式翻译效果
- 适合追求低延迟高准确率的实时翻译系统开发者
同时语音翻译(SimulST)系统在接收音频的同时实时生成译文,需在翻译质量与延迟之间取得平衡。本文提出一种新策略:仅当等待能带来信息增益时才继续等待。基于此,我们设计了正则化熵信息适应(REINA)损失函数,利用现有非流式翻译模型训练自适应决策策略。该方法基于信息论推导,可推动延迟/质量权衡的帕累托前沿超越此前工作。我们在法语、西班牙语、德语与英语之间的双向翻译任务上训练模型,仅使用开源或合成数据,便达到与模型规模相当的最先进流式翻译性能。我们还引入流式效率指标,量化显示相较于以往方法,REINA在标准化后的非流式基准BLEU得分下,延迟-质量权衡提升最高达21%。
原文摘要 · Abstract (English)
Simultaneous Speech Translation (SimulST) systems stream in audio while simultaneously emitting translated text or speech. Such systems face the significant challenge of balancing translation quality and latency. We introduce a strategy to optimize this tradeoff: wait for more input only if you gain information by doing so. Based on this strategy, we present Regularized Entropy INformation Adaptation (REINA), a novel loss to train an adaptive policy using an existing non-streaming translation model. We derive REINA from information theory principles and show that REINA helps push the reported Pareto frontier of the latency/quality tradeoff over prior works. Utilizing REINA, we train a SimulST model on French, Spanish and German, both from and into English. Training on only open source or synthetically generated data, we achieve state-of-the-art (SOTA) streaming results for models of comparable size. We also introduce a metric for streaming efficiency, quantitatively showing REINA improves the latency/quality trade-off by as much as 21% compared to prior approaches, normalized against non-streaming baseline BLEU scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。