arXiv:2505.19179cs.SDeess.AS2025-05中稿 · InterSpeech 2025被引 17

提升语音大模型对专有名词的识别准确率,支持超大规模偏置词库。

BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM

  • 通过声学与偏置对比学习检索语义相关候选词
  • 2000个偏置词下实现2.8%/7.1%的最优偏置词错误率
  • 支持20万条偏置词,延迟仅20ms,适合工业级语音系统

尽管语音大语言模型(SpeechLLMs)已推动标准自动语音识别(ASR)的发展,但在大规模场景下对专有名词和生僻词进行上下文偏置仍具挑战。为此,我们提出BR-ASR:一种基于双创新的大规模上下文偏置检索框架(支持最多20万条词条)。其一,采用语音与偏置对比学习机制,高效检索语义相关候选词;其二,引入动态课程学习策略,缓解同音词混淆问题。该框架无需微调即可无缝集成至多种ASR系统。在LibriSpeech test-clean/-other上的实验表明,使用2000个偏置词时,取得2.8%/7.1%的最优偏置词错误率(B-WER),相较以往方法相对提升45%。此外,当偏置列表扩展至20万条时,传统方法普遍失效,而BR-ASR仅带来0.3%/2.9%的绝对词错误率(WER)/B-WER增加,剪枝率达99.99%,单次查询延迟仅为20ms。

原文摘要 · Abstract (English)

While speech large language models (SpeechLLMs) have advanced standard automatic speech recognition (ASR), contextual biasing for named entities and rare words remains challenging, especially at scale. To address this, we propose BR-ASR: a Bias Retrieval framework for large-scale contextual biasing (up to 200k entries) via two innovations: (1) speech-and-bias contrastive learning to retrieve semantically relevant candidates; (2) dynamic curriculum learning that mitigates homophone confusion which negatively impacts the final performance. The is a general framework that allows seamless integration of the retrieved candidates into diverse ASR systems without fine-tuning. Experiments on LibriSpeech test-clean/-other achieve state-of-the-art (SOTA) biased word error rates (B-WER) of 2.8%/7.1% with 2000 bias words, delivering 45% relative improvement over prior methods. BR-ASR also demonstrates high scalability: when expanding the bias list to 200k where traditional methods generally fail, it induces only 0.3 / 2.9% absolute WER / B-WER degradation with a 99.99% pruning rate and only 20ms latency per query on test-other.

语音识别偏置检索大模型应用高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。