arXiv:2509.19567cs.CLeess.AS2025-09EMNLP被引 2

用检索增强生成自动发现语音识别上下文,提升罕见词识别准确率。

Retrieval Augmented Generation based context discovery for ASR

  • 基于嵌入的检索方法自动挖掘语音识别上下文
  • 在多个数据集上将错误率降低最多17%
  • 适合需要提升低频词识别的语音系统开发者

本文研究了检索增强生成在上下文感知语音识别系统中用于自动上下文发现的高效策略,旨在提升面对罕见词或未登录词时的转录准确率。然而,自动识别合适上下文仍是开放挑战。本文提出一种基于嵌入的高效检索方法,用于语音识别中的自动上下文发现。为验证其有效性,还评估了两种基于大语言模型(LLM)的替代方案:(1) 通过提示生成上下文;(2) 使用大语言模型对识别后文本进行修正。在TED-LIUMv3、Earnings21和SPGISpeech数据集上的实验表明,所提方法相比无上下文情况,将词错误率(WER)降低最多17%(相对减少),而理想上下文可带来最高24.1%的降低。

原文摘要 · Abstract (English)

This work investigates retrieval augmented generation as an efficient strategy for automatic context discovery in context-aware Automatic Speech Recognition (ASR) system, in order to improve transcription accuracy in the presence of rare or out-of-vocabulary terms. However, identifying the right context automatically remains an open challenge. This work proposes an efficient embedding-based retrieval approach for automatic context discovery in ASR. To contextualize its effectiveness, two alternatives based on large language models (LLMs) are also evaluated: (1) large language model (LLM)-based context generation via prompting, and (2) post-recognition transcript correction using LLMs. Experiments on the TED-LIUMv3, Earnings21 and SPGISpeech demonstrate that the proposed approach reduces WER by up to 17% (percentage difference) relative to using no-context, while the oracle context results in a reduction of up to 24.1%.

语音识别检索增强上下文发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。