arXiv:2505.14600eess.AScs.SD2025-05中稿 · Interspeech 2025被引 3

AdaKWS通过测试时自适应提升语音关键词识别鲁棒性

AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation

  • 基于预测熵最小化筛选可靠样本,动态调整批量归一化统计量
  • 引入伪关键词一致性机制,有效识别关键特征避免噪声过拟合
  • 在真实噪声场景下显著优于现有方法,适合边缘设备部署

语音关键词识别(KWS)旨在从音频中识别关键词,广泛应用于边缘设备。现有轻量级KWS系统虽注重模型效率,但在未见环境或噪声背景下性能下降。测试时自适应(TTA)可在不依赖原始训练数据的情况下使模型适应测试样本。本文提出AdaKWS,据我们所知是首个用于鲁棒KWS的TTA方法。首先,通过最小化预测熵筛选可靠样本,并动态调整每批次的归一化统计量以优化模型置信度;其次,引入伪关键词一致性(PKC)机制,在不过拟合噪声的前提下识别关键可靠特征。实验表明,AdaKWS在高斯噪声和真实场景噪声下均优于其他方法。代码将于后续发布。

原文摘要 · Abstract (English)

Spoken keyword spotting (KWS) aims to identify keywords in audio for wide applications, especially on edge devices. Current small-footprint KWS systems focus on efficient model designs. However, their inference performance can decline in unseen environments or noisy backgrounds. Test-time adaptation (TTA) helps models adapt to test samples without needing the original training data. In this study, we present AdaKWS, the first TTA method for robust KWS to the best of our knowledge. Specifically, 1) We initially optimize the model's confidence by selecting reliable samples based on prediction entropy minimization and adjusting the normalization statistics in each batch. 2) We introduce pseudo-keyword consistency (PKC) to identify critical, reliable features without overfitting to noise. Our experiments show that AdaKWS outperforms other methods across various conditions, including Gaussian noise and real-scenario noises. The code will be released in due course.

关键词识别测试时自适应边缘计算噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。