提升语音增强语言模型在复杂场景下的实体识别准确率
COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

- 通过对比正则化将语音表示映射到判别空间,量化音频与候选实体匹配度
- 在不同规模偏置列表下均优于基线,在多实体共现时表现更稳定
- 适合需要精准识别专业术语的语音识别场景,如医疗、法律领域
上下文偏置旨在将外部知识融入自动语音识别(ASR)系统,以准确识别领域特定实体。本文提出COALA(Contextualized ASR Leveraging Biasing Scoring),一种鲁棒框架,用于提升复杂多实体场景下语音增强语言模型(SLMs)的性能。考虑到SLMs固有的上下文窗口限制,从大规模偏置列表中识别相关目标实体至关重要。为此,COALA将SLM隐层表示映射至专用判别空间,量化音频片段与候选实体之间的匹配强度。此外,针对先前研究在处理多目标话语(多个罕见词共现)时出现的训练崩溃问题,本方法有效缓解。在LibriSpeech基准上的实验表明,COALA在不同规模偏置列表下均持续实现更优的上下文偏置性能。
原文摘要 · Abstract (English)
Contextual biasing seeks to integrate external knowledge into automatic speech recognition (ASR) systems to accurately recognize domain-specific entities. In this paper, we propose COALA (Contextualized ASR Leveraging Biasing Scoring), a robust framework designed to enhance speech-augmented language models (SLMs) in complex multi-entity scenarios. Considering the inherent context-window limitations of SLMs, identifying relevant target entities from a large-scale biasing list is crucial for effective recognition. To this end, COALA maps SLM latent representations into a specialized discriminative space to quantify the matching intensity between audio segments and candidate entities. Furthermore, we address the training collapse in prior study when handling multi-target utterances-where multiple rare words co-occur. Experimental results on the LibriSpeech benchmark demonstrate that COALA consistently achieves superior contextual biasing performance across various biasing list scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。