arXiv:2503.14185cs.CLcs.SD2025-03被引 4

让语音翻译模型动态调整声学状态,提升跨模态翻译效果

AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation

  • decoder根据自身状态动态调整编码器的声学表示
  • 在两个数据集上显著超越现有最优模型
  • 适合研究语音翻译与跨模态交互的学者

端到端语音翻译中,编码器学习的声学表征通常对解码器固定不变,难以应对语音翻译中的跨模态和跨语言挑战。本文提出一种自适应语音转文本翻译模型,使解码器能够动态调整声学状态。通过将声学状态与目标词嵌入序列拼接,并输入解码器后续模块,引入语音-文本混合注意力子层替代传统交叉注意力网络,以建模声学状态与目标隐藏状态间的深层交互。在两个常用数据集上的实验结果表明,该方法显著优于当前最先进的神经语音翻译模型。

原文摘要 · Abstract (English)

In end-to-end speech translation, acoustic representations learned by the encoder are usually fixed and static, from the perspective of the decoder, which is not desirable for dealing with the cross-modal and cross-lingual challenge in speech translation. In this paper, we show the benefits of varying acoustic states according to decoder hidden states and propose an adaptive speech-to-text translation model that is able to dynamically adapt acoustic states in the decoder. We concatenate the acoustic state and target word embedding sequence and feed the concatenated sequence into subsequent blocks in the decoder. In order to model the deep interaction between acoustic states and target hidden states, a speech-text mixed attention sublayer is introduced to replace the conventional cross-attention network. Experiment results on two widely-used datasets show that the proposed method significantly outperforms state-of-the-art neural speech translation models.

语音翻译动态适配注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。