arXiv:2409.11889cs.SDeess.AS2024-09中稿 · ICASSP 2025, oral被引 8

用多阶段检索增强,让Whisper更好识别方言。

M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper

  • 预处理用句级上下文学习,后处理用词级最近邻检索。
  • 在中文和方言数据集上准确率显著提升,无需修改模型参数。
  • 适合低资源场景下提升语音识别精度,尤其对方言有效。

当前最先进的模型如OpenAI的Whisper在多语言自动语音识别(ASR)中表现优异,但仍难以准确识别多样化的子方言。本文提出M2R-Whisper,一种新型多阶段、多尺度检索增强方法,旨在提升低资源环境下的ASR性能。基于上下文学习(ICL)与检索增强技术,该方法在预处理阶段采用句级ICL以利用上下文信息,在后处理阶段引入词级k-近邻(kNN)检索,进一步优化最终输出分布。通过句级与词级检索策略的协同作用,M2R-Whisper有效缓解了多种识别错误。在普通话及子方言数据集(包括AISHELL-1和KeSpeech)上的实验表明,该方法在不更新任何模型参数的情况下实现了显著的ASR准确率提升。

原文摘要 · Abstract (English)

State-of-the-art models like OpenAI's Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-whisper, a novel multi-stage and multi-scale retrieval augmentation approach designed to enhance ASR performance in low-resource settings. Building on the principles of in-context learning (ICL) and retrieval-augmented techniques, our method employs sentence-level ICL in the pre-processing stage to harness contextual information, while integrating token-level k-Nearest Neighbors (kNN) retrieval as a post-processing step to further refine the final output distribution. By synergistically combining sentence-level and token-level retrieval strategies, M2R-whisper effectively mitigates various types of recognition errors. Experiments conducted on Mandarin and subdialect datasets, including AISHELL-1 and KeSpeech, demonstrate substantial improvements in ASR accuracy, all achieved without any parameter updates.

语音识别方言识别检索增强Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。