用大模型提升语音识别中罕见词的零样本表现。
Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration
- 将大语言模型与语音识别系统结合,改进罕见词识别。
- 在19万小时数据上,罕见词错误率显著降低。
- 适合关注大模型融合语音识别的研究者。
本研究探讨将大语言模型(LLM)集成到自动语音识别(ASR)系统中,以提升罕见词识别性能。基于主要来自YouTube的19万小时数据集,经Whisper V3伪标签预处理,实验表明,该LLM-ASR架构在零样本罕见词识别任务中优于传统Zipformer-Transducer模型。分析显示,LLM显著改善罕见词错误率(R-WER),而语音编码器主导整体转录性能(正字法词错误率O-WER与归一化词错误率N-WER)。通过大量消融实验,揭示适配器集成对对齐语音编码器输出与LLM语言能力的关键作用。同时强调高质量标注数据对达到最优性能的重要性。研究为基于大模型的语音识别架构协同提供了重要洞见。
原文摘要 · Abstract (English)
In this study, we investigate the integration of a large language model (LLM) with an automatic speech recognition (ASR) system, specifically focusing on enhancing rare word recognition performance. Using a 190,000-hour dataset primarily sourced from YouTube, pre-processed with Whisper V3 pseudo-labeling, we demonstrate that the LLM-ASR architecture outperforms traditional Zipformer-Transducer models in the zero-shot rare word recognition task, after training on a large dataset. Our analysis reveals that the LLM contributes significantly to improvements in rare word error rate (R-WER), while the speech encoder primarily determines overall transcription performance (Orthographic Word Error Rate, O-WER, and Normalized Word Error Rate, N-WER). Through extensive ablation studies, we highlight the importance of adapter integration in aligning speech encoder outputs with the LLM's linguistic capabilities. Furthermore, we emphasize the critical role of high-quality labeled data in achieving optimal performance. These findings provide valuable insights into the synergy between LLM-based ASR architectures, paving the way for future advancements in large-scale LLM-based speech recognition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。