针对长文本语音识别,动态筛选关键信息提升准确率。
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
- 分两阶段用语音驱动机制压缩并保留重要上下文
- 在两个数据集上分别达7.71%和1.12%的词错误率
- 特别适合会议等需领域知识的语音识别场景
自动语音识别(ASR)系统在常规场景下表现优异,但在需要领域知识的长上下文场景(如会议演讲)中难以有效利用上下文信息。主要原因是模型上下文窗口受限及上下文中相关信号稀疏。为此,我们提出SAP²方法,通过两阶段动态剪枝与整合相关上下文关键词。每个阶段均采用提出的语音驱动注意力池化机制,高效压缩上下文嵌入的同时保留语音显著信息。实验表明,SAP²在SlideSpeech和LibriSpeech数据集上达到当前最优性能,词错误率(WER)分别为7.71%和1.12%。在SlideSpeech上,相比非上下文基线,偏见关键词错误率(B-WER)降低41.1%。SAP²还表现出强可扩展性,在大量上下文输入下仍保持稳定性能。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems have achieved remarkable performance in common conditions but often struggle to leverage long-context information in contextualized scenarios that require domain-specific knowledge, such as conference presentations. This challenge arises primarily due to constrained model context windows and the sparsity of relevant information within extensive contextual noise. To solve this, we propose the SAP$^{2}$ method, a novel framework that dynamically prunes and integrates relevant contextual keywords in two stages. Specifically, each stage leverages our proposed Speech-Driven Attention-based Pooling mechanism, enabling efficient compression of context embeddings while preserving speech-salient information. Experimental results demonstrate state-of-the-art performance of SAP$^{2}$ on the SlideSpeech and LibriSpeech datasets, achieving word error rates (WER) of 7.71% and 1.12%, respectively. On SlideSpeech, our method notably reduces biased keyword error rates (B-WER) by 41.1% compared to non-contextual baselines. SAP$^{2}$ also exhibits robust scalability, consistently maintaining performance under extensive contextual input conditions on both datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。