arXiv:2410.03298eess.AS2024-10被引 6

无需文本中间步骤,直接生成语音翻译,低延迟效果更优。

Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

  • 用声学令牌直接输出语音,跳过文本转换环节。
  • 在多个语种对上达到最优的BLEU与低延迟表现。
  • 适合实时语音翻译场景,尤其关注低延迟应用。

级联式语音到语音翻译系统常因模块累积推理延迟导致误差传播和高延迟问题。本文提出一种基于变换器的语音翻译模型,以低延迟流式方式输出离散语音令牌。该方法无需先生成文本,再经机器翻译(MT)和文本到语音(TTS)转换。生成的语音令牌可借助声学语言模型(LM)获取声学令牌,并通过音频编解码模型还原波形,实现低延迟语音合成。实验结果表明,该方法在多个语言对上优于现有方法,在CVSS-C数据集上的基准测试中,于BLEU、平均延迟及BLASER 2.0评分上均达当前最佳水平。

原文摘要 · Abstract (English)

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech translation model that outputs discrete speech tokens in a low-latency streaming fashion. This approach eliminates the need for generating text output first, followed by machine translation (MT) and text-to-speech (TTS) systems. The produced speech tokens can be directly used to generate a speech signal with low latency by utilizing an acoustic language model (LM) to obtain acoustic tokens and an audio codec model to retrieve the waveform. Experimental results show that the proposed method outperforms other existing approaches and achieves state-of-the-art results for streaming translation in terms of BLEU, average latency, and BLASER 2.0 scores for multiple language pairs using the CVSS-C dataset as a benchmark.

语音翻译低延迟流式处理无文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。