arXiv:2411.07387cs.CLeess.AS2024-11

让语音翻译时长精准匹配原文,提升跨语种同步性。

Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages

  • 在解码时引入时间信息,同步控制语音与停顿的持续时间。
  • 在中文转英文测试集上实现0.92语音重叠率和8.9 BLEU值。
  • 适合对时长同步要求高的语音翻译场景,如实时字幕、双语播报。

端到端语音翻译(ST)近年来受到广泛关注,其目标是将源语言语音直接转化为目标语言文本。许多应用场景要求翻译结果的时长严格匹配源音频长度,包括语音和停顿段落。现有方法通常通过控制机器翻译生成的词数或字符数来近似源句长度,但未考虑不同语言中语音与停顿的等时性差异。为此,本文改进了序列到序列ST模型中的时长对齐组件,提出基于等时性的语音翻译控制方法:在翻译过程中联合预测语音和停顿的持续时间,并向解码器提供时间信息,使其在生成时跟踪剩余语音与停顿时长。在CoVoST 2的中文-英文测试集上的评估表明,该方法实现了0.92的语音重叠率和8.9的BLEU值,仅比基线模型低1.4 BLEU。

原文摘要 · Abstract (English)

End-to-end speech translation (ST), which translates source language speech directly into target language text, has garnered significant attention in recent years. Many ST applications require strict length control to ensure that the translation duration matches the length of the source audio, including both speech and pause segments. Previous methods often controlled the number of words or characters generated by the Machine Translation model to approximate the source sentence's length without considering the isochrony of pauses and speech segments, as duration can vary between languages. To address this, we present improvements to the duration alignment component of our sequence-to-sequence ST model. Our method controls translation length by predicting the duration of speech and pauses in conjunction with the translation process. This is achieved by providing timing information to the decoder, ensuring it tracks the remaining duration for speech and pauses while generating the translation. The evaluation on the Zh-En test set of CoVoST 2, demonstrates that the proposed Isochrony-Controlled ST achieves 0.92 speech overlap and 8.9 BLEU, which has only a 1.4 BLEU drop compared to the ST baseline.

语音翻译时长控制等时性跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。