仅用一次学习即可让语音关键词识别系统自适应降噪,适合嵌入式设备部署。
Adaptive Noise Resilient Keyword Spotting Using One-Shot Learning
- 通过一次学习和一个训练轮次,动态调整预训练模型以适应噪声环境。
- 在信噪比低至-3 dB时,准确率提升4.9%到46.0%,尤其在18 dB以下表现突出。
- 方法轻量高效,适合内存与算力受限的智能设备,支持实时更新关键词。
关键词识别(KWS)是智能设备实现高效语音交互的核心技术。然而,传统部署于嵌入式设备的KWS系统在真实环境下常因性能下降而受限。抗干扰型KWS系统通过动态适应机制解决此问题,支持关键词增删、用户适配及噪声鲁棒性提升。但资源受限设备上实现低延迟、独立运行的抗干扰KWS仍具挑战。本研究提出一种低计算量的连续噪声适应方法,仅需1次学习和1个训练轮次,即可对预训练神经网络进行微调。实验采用两个预训练模型,在三种真实噪声源下,于24至-3 dB的信噪比范围内评估。适应后的模型在所有场景中均优于原始模型,尤其在信噪比≤18 dB时,准确率提升达4.9%至46.0%。结果表明该方法高效可靠,适合在资源受限设备上部署。
原文摘要 · Abstract (English)
Keyword spotting (KWS) is a key component of smart devices, enabling efficient and intuitive audio interaction. However, standard KWS systems deployed on embedded devices often suffer performance degradation under real-world operating conditions. Resilient KWS systems address this issue by enabling dynamic adaptation, with applications such as adding or replacing keywords, adjusting to specific users, and improving noise robustness. However, deploying resilient, standalone KWS systems with low latency on resource-constrained devices remains challenging due to limited memory and computational resources. This study proposes a low computational approach for continuous noise adaptation of pretrained neural networks used for KWS classification, requiring only 1-shot learning and one epoch. The proposed method was assessed using two pretrained models and three real-world noise sources at signal-to-noise ratios (SNRs) ranging from 24 to -3 dB. The adapted models consistently outperformed the pretrained models across all scenarios, especially at SNR $\leq$ 18 dB, achieving accuracy improvements of 4.9% to 46.0%. These results highlight the efficacy of the proposed methodology while being lightweight enough for deployment on resource-constrained devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。