arXiv:2505.23077cs.SDcs.CL2025-05中稿 · interspeech 2025

动态词汇预测让语音识别更准,尤其擅长完整输出上下文短语。

Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation

  • 用动态词汇预测和激活机制,按短语级别增强上下文信息。
  • 在Librispeech和Wenetspeech上,关键词错误率降低超23%,短语错误率降75%。
  • 适合需要高精度上下文理解的语音识别场景,如医疗、法律录音转写。

深度上下文引导可提升语音识别性能,但现有方法将上下文短语拆分为独立子词单元处理,可能破坏短语完整性,导致准确率下降。本文提出一种基于编码器的短语级上下文感知语音识别方法,引入动态词汇预测与激活机制。通过架构优化并结合偏置损失函数,从帧级输出扩展出短语级预测结果。同时设计置信度激活解码策略,在确保完整输出上下文短语的同时抑制错误干扰。在Librispeech和Wenetspeech数据集上的实验表明,相比基线模型,相对单词错误率分别降低28.31%和23.49%,上下文短语的相对错误率分别下降72.04%和75.69%。

原文摘要 · Abstract (English)

Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromising contextual phrase integrity, leading to accuracy reduction. In this paper, we propose an encoder-based phrase-level contextualized ASR method that leverages dynamic vocabulary prediction and activation. We introduce architectural optimizations and integrate a bias loss to extend phrase-level predictions based on frame-level outputs. We also introduce a confidence-activated decoding method that ensures the complete output of contextual phrases while suppressing incorrect bias. Experiments on Librispeech and Wenetspeech datasets demonstrate that our approach achieves relative WER reductions of 28.31% and 23.49% compared to baseline, with the WER on contextual phrases decreasing relatively by 72.04% and 75.69%.

语音识别上下文建模动态词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。