用预训练语音模型和大语言模型实现语音识别与翻译的一体化,性能超越现有方案。
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
- 融合预训练语音编码器与大语言模型,统一处理语音识别与翻译
- 英德语对下比SeamlessM4T提升8%的COMETDA22得分,媲美级联系统
- 适合追求端到端高效语音多语言处理的研究者与开发者
语音翻译(ST)是将一种语言的语音信号转换为另一种语言对应文本的机器翻译任务,传统方法采用级联架构,近年兴起端到端方案。本文探索将预训练语音编码器与大语言模型(LLMs)结合的端到端架构,同时完成自动语音识别(ASR)与语音翻译。在英德语对上的实验表明,最优模型不仅在翻译效果上优于SeamlessM4T这一大型基础端到端多模态翻译模型,且性能可媲美由Whisper与NLLB组成的级联系统,在COMETDA22指标上最高提升达8%。
原文摘要 · Abstract (English)
Speech Translation (ST) is a machine translation task that involves converting speech signals from one language to the corresponding text in another language; this task has two different approaches, namely the traditional cascade and the more recent end-to-end. This paper explores a combined end-to-end architecture of pre-trained speech encoders and Large Language Models (LLMs) for performing both Automatic Speech Recognition (ASR) and ST simultaneously. Experiments with the English-to-German language pair show that our best model not only can achieve better translation results than SeamlessM4T, a large foundational end-to-end, multi-modal translation model, but can also match the performance of a cascaded system with Whisper and NLLB, with up to a score gain of 8% in $\text{COMET}^{\text{DA}}_{22}$ metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。