用语音检索增强大模型,提升方言识别准确率
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
- 构建细粒度语音数据池,通过语音检索增强模型上下文
- 在普通话和方言数据集上准确率显著提升,尤其改善口音识别
- 适合需要高鲁棒性语音识别的场景,如多语种语音系统
将语音信息融入大语言模型(LLM)的最新进展显著提升了自动语音识别(ASR)的准确率。然而,现有方法受限于语音编码器在不同声学条件(如口音)下的表现。为此,我们提出LA-RAG,一种面向基于大模型的ASR的新颖检索增强生成(RAG)范式。LA-RAG利用细粒度的词元级语音数据存储和语音到语音的检索机制,通过大模型的上下文学习(ICL)能力增强ASR性能。在普通话及多种中文方言数据集上的实验表明,相比现有方法,该方法在提升识别准确率方面具有显著优势,尤其在处理口音差异方面表现突出。
原文摘要 · Abstract (English)
Recent advancements in integrating speech information into large language models (LLMs) have significantly improved automatic speech recognition (ASR) accuracy. However, existing methods often constrained by the capabilities of the speech encoders under varied acoustic conditions, such as accents. To address this, we propose LA-RAG, a novel Retrieval-Augmented Generation (RAG) paradigm for LLM-based ASR. LA-RAG leverages fine-grained token-level speech datastores and a speech-to-speech retrieval mechanism to enhance ASR accuracy via LLM in-context learning (ICL) capabilities. Experiments on Mandarin and various Chinese dialect datasets demonstrate significant improvements in ASR accuracy compared to existing methods, validating the effectiveness of our approach, especially in handling accent variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。