提出噪声感知的关键词检测框架,显著提升低信噪比下的识别准确率。
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
- 在CTC框架中引入噪声建模的自环与旁路弧,增强对噪声的鲁棒性。
- 在极低信噪比下,误报率降低37%,性能优于现有端到端系统。
- 适合部署在资源受限设备上的语音唤醒系统,尤其适用于嘈杂环境。
近年来,基于连接时序分类的轻量级关键词检测(CTC-KWS)系统受到广泛关注。这类系统通常部署在计算资源受限的设备上,在复杂声学环境下易因模型容量不足导致过拟合及关键词与背景噪声混淆,引发高误报。为此,本文提出一种噪声感知的CTC关键词检测框架(NTC-KWS),旨在提升低信噪比场景下的鲁棒性。该方法基于加权有限状态转换器(WFST)图,在训练与解码过程中引入两类额外的噪声建模通路:用于处理噪声插入错误的自环弧,以及应对过度噪声遮蔽与干扰的旁路弧。在干净与嘈杂的Hey Snips数据集上实验表明,NTC-KWS在多种声学条件下均超越当前最优端到端系统和CTC-KWS基线,尤其在低信噪比场景中表现突出。
原文摘要 · Abstract (English)
In recent years, there has been a growing interest in designing small-footprint yet effective Connectionist Temporal Classification based keyword spotting (CTC-KWS) systems. They are typically deployed on low-resource computing platforms, where limitations on model size and computational capacity create bottlenecks under complicated acoustic scenarios. Such constraints often result in overfitting and confusion between keywords and background noise, leading to high false alarms. To address these issues, we propose a noise-aware CTC-based KWS (NTC-KWS) framework designed to enhance model robustness in noisy environments, particularly under extremely low signal-to-noise ratios. Our approach introduces two additional noise-modeling wildcard arcs into the training and decoding processes based on weighted finite state transducer (WFST) graphs: self-loop arcs to address noise insertion errors and bypass arcs to handle masking and interference caused by excessive noise. Experiments on clean and noisy Hey Snips show that NTC-KWS outperforms state-of-the-art (SOTA) end-to-end systems and CTC-KWS baselines across various acoustic conditions, with particularly strong performance in low SNR scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。