arXiv:2409.16654eess.AScs.CL2024-09被引 6

用多模态大模型重打语音识别结果,效果提升超20%

Speech Recognition Rescoring with Large Speech-Text Foundation Models

  • 用语音-文本基础模型做语音识别后处理重打分
  • 相比Whisper大模型提升最高20%,文本模型提升15%
  • 适合需要高精度语音识别的场景,如医疗、法律

大型语言模型(LLM)通过大量文本数据展现出强大的语言理解能力。自动语音识别(ASR)系统常受限于标注语音数据量,可通过使用LLM进行二次重打分来提升性能。近期,多模态大语言模型,尤其是语音与文本基础模型,在口语理解方面表现优异。这类模型利用语音和文本模态中的大量无标签与有标签数据建模人类语言。本文提出新颖方法,利用多模态LLM进行ASR重打分,并探索判别式训练以进一步提升性能。实验表明,语音-文本基础模型中的跨模态知识迁移可显著提升重打分效果。在多个数据集上,相比Whisper large ASR系统,相对改进最高达20%;相比仅文本的LLM,相对改进最高达15%。

原文摘要 · Abstract (English)

Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech data and benefit from a second pass rescoring using LLM. Recently multi-modal large language models, particularly speech and text foundational models have demonstrated strong spoken language understanding. Speech-Text foundational models leverage large amounts of unlabelled and labelled data both in speech and text modalities to model human language. In this work, we propose novel techniques to use multi-modal LLM for ASR rescoring. We also explore discriminative training to further improve the foundational model rescoring performance. We demonstrate cross-modal knowledge transfer in speech-text LLM can benefit rescoring. Our experiments demonstrate up-to 20% relative improvements over Whisper large ASR and up-to 15% relative improvements over text-only LLM.

语音识别多模态模型重打分大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。