arXiv:2510.10003cs.CLcs.SD2025-10被引 1

通过多标记预测提升语音到语音翻译的语义密度与质量

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction

  • 在中间层引入多标记预测损失,提前增强隐藏表示
  • 相较单标记预测,多标记预测显著提升翻译质量
  • 适合关注语音翻译模型优化的研究者

当前直接语音到语音翻译方法主要使用语音标记作为中间表示。然而,单个语音标记语义密度不足,通常需要多个标记才能表达一个完整语义单元。为此,本文在语音到单元翻译(S2UT)模型中引入多标记预测(MTP)损失,使模型在每个位置预测多个后续标记,从而捕获更完整的语义并提升每位置的信息密度。早期的MTP实现将损失应用于最后层,虽改善输出表示,但信息丰富化过晚。我们假设将信息丰富化过程提前至中间层可实现更早、更有效的隐藏表示增强。因此,提出MTP-S2UT损失,将其应用于计算CTC损失的隐藏表示层。实验表明,所有MTP损失变体均持续提升S2UT翻译质量,其中MTP-S2UT表现最佳。

原文摘要 · Abstract (English)

Current direct speech-to-speech translation methods predominantly employ speech tokens as intermediate representations. However, a single speech token is not dense in semantics, so we generally need multiple tokens to express a complete semantic unit. To address this limitation, we introduce multi-token prediction (MTP) loss into speech-to-unit translation (S2UT) models, enabling models to predict multiple subsequent tokens at each position, thereby capturing more complete semantics and enhancing information density per position. Initial MTP implementations apply the loss at the final layer, which improves output representation but initiates information enrichment too late. We hypothesize that advancing the information enrichment process to intermediate layers can achieve earlier and more effective enhancement of hidden representation. Consequently, we propose MTP-S2UT loss, applying MTP loss to hidden representation where CTC loss is computed. Experiments demonstrate that all MTP loss variants consistently improve the quality of S2UT translation, with MTP-S2UT achieving the best performance.

语音翻译多标记预测S2UT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。