arXiv:2412.00721cs.AIcs.CL2024-12被引 5

对比LLM与Whisper在低资源和中英混用场景下的语音识别表现

A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario

  • 用大语言模型构建语音识别系统,结合声学编码器提升性能
  • 低资源场景下比Whisper相对提升12.8%,但混用场景下Whisper更优
  • 为资源匮乏语言和混用语境的语音识别提供实证参考

大语言模型(LLMs)在多种自然语言任务中表现出色,其与语音编码器的结合正迅速成为自动语音识别(ASR)领域的主流趋势。以往研究多集中于英语和中文场景下的应用,而对低资源环境下语音识别潜力的探索仍不足。本文旨在评估LLM在低资源ASR及普通话-英语混用场景中的表现,并与Whisper模型进行对比。大量实验表明,在低资源场景下,基于LLM的ASR相比Whisper相对提升12.8%;而在中英混用场景中,Whisper表现更佳。本研究为低资源场景下的语音识别提供了重要参考。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) field. Previous works mainly concentrated on leveraging LLMs for speech recognition in English and Chinese. However, their potential for addressing speech recognition challenges in low resource settings remains underexplored. Hence, in this work, we aim to explore the capability of LLMs in low resource ASR and Mandarin-English code switching ASR. We also evaluate and compare the recognition performance of LLM-based ASR systems against Whisper model. Extensive experiments demonstrate that LLM-based ASR yields a relative gain of 12.8\% over the Whisper model in low resource ASR while Whisper performs better in Mandarin-English code switching ASR. We hope that this study could shed light on ASR for low resource scenarios.

语音识别低资源混用语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。