arXiv:2412.15415eess.AScs.CL2024-12被引 2

端到端联合语音识别与翻译,实时低延迟。

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition

  • 基于快慢级联编码器的联合建模,支持流式处理。
  • 在双语对话场景中,翻译准确率更高且延迟更低。
  • 适合智能眼镜等实时交互设备使用。

我们提出一种联合语音翻译与识别(JSTAR)模型,采用快-慢级联编码器架构,实现端到端的流式自动语音识别(ASR)与语音翻译(ST)。该模型基于转换器结构,采用多目标训练策略,同时优化ASR与ST目标,从而生成高质量的实时结果。在智能眼镜支持的双语对话场景中,模型还被训练以区分佩戴者与对话者的声音方向。研究了多种预训练策略,首次基于转换器的流式机器翻译(MT)模型用于参数初始化。实验表明,相较于强基准的级联式ST模型,JSTAR在BLEU分数和延迟方面均表现更优。

原文摘要 · Abstract (English)

We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recognition (ASR) and speech translation (ST). The model is transducer-based and uses a multi-objective training strategy that optimizes both ASR and ST objectives simultaneously. This allows JSTAR to produce high-quality streaming ASR and ST results. We apply JSTAR in a bilingual conversational speech setting with smart-glasses, where the model is also trained to distinguish speech from different directions corresponding to the wearer and a conversational partner. Different model pre-training strategies are studied to further improve results, including training of a transducer-based streaming machine translation (MT) model for the first time and applying it for parameter initialization of JSTAR. We demonstrate superior performances of JSTAR compared to a strong cascaded ST model in both BLEU scores and latency.

语音翻译端到端流式处理智能眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。