arXiv:2412.12635eess.AScs.SD2024-12中稿 · ICASSP2025被引 7

提出跨层判别一致性机制,提升流式关键词检测准确率

Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency

  • 设计轻量级流式解码算法,可任意位置检测关键词起点
  • 跨层判别一致性使正负样本区分更精准,召回率提升6.8%
  • 适合对低误报率要求高的实时语音交互系统

连接时序分类(CTC)作为非自回归训练准则,广泛应用于在线关键词检测(KWS)。然而,现有基于CTC的KWS解码策略或依赖自动语音识别(ASR),因未针对关键词优化而性能受限;或使用专用解码图,实现复杂且维护困难。本文提出一种面向CTC-KWS的流式解码算法,引入跨层判别一致性(CDC)机制。该算法能任意位置检测关键词起始点,并利用多层间判别一致性信息增强正负样本区分能力。在干净与嘈杂的Hey Snips数据集上实验表明,所提方法优于基于ASR和图结构的基线。引入CDC后,平均召回率提升6.8%,漏检率相对降低46.3%,误报率仅为每小时0.05次。

原文摘要 · Abstract (English)

Connectionist Temporal Classification (CTC), a non-autoregressive training criterion, is widely used in online keyword spotting (KWS). However, existing CTC-based KWS decoding strategies either rely on Automatic Speech Recognition (ASR), which performs suboptimally due to its broad search over the acoustic space without keyword-specific optimization, or on KWS-specific decoding graphs, which are complex to implement and maintain. In this work, we propose a streaming decoding algorithm enhanced by Cross-layer Discrimination Consistency (CDC), tailored for CTC-based KWS. Specifically, we introduce a streamlined yet effective decoding algorithm capable of detecting the start of the keyword at any arbitrary position. Furthermore, we leverage discrimination consistency information across layers to better differentiate between positive and false alarm samples. Our experiments on both clean and noisy Hey Snips datasets show that the proposed streaming decoding strategy outperforms ASR-based and graph-based KWS baselines. The CDC-boosted decoding further improves performance, yielding an average absolute recall improvement of 6.8% and a 46.3% relative reduction in the miss rate compared to the graph-based KWS baseline, with a very low false alarm rate of 0.05 per hour.

关键词检测流式处理深度学习语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。