arXiv:2510.15432eess.AScs.SD2025-10

通过量化嵌入改善噪声下少样本关键词检测的阈值选择

Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments

  • 在嵌入层对特征进行量化,利用量化误差归一化
  • 在模拟高频无线电信道上提升检测准确率
  • 适合部署在噪声环境下的少样本关键词识别系统

关键词检测(KWS)系统需对连续检测分数进行阈值判断。传统方法依赖验证集优化阈值,但在未知或噪声环境下、特别是在少样本设置中常表现不佳。本文研究基于动态时间规整(DTW)的模板式开集少样本KWS在噪声语音中的阈值估计问题。为缓解次优阈值导致的性能下降,提出一种在嵌入层操作的分数校准方法:通过量化学习到的表示,并在DTW评分与阈值判断前应用基于量化误差的归一化。在包含模拟高频无线电通道的KWS-DailyTalk数据集上的实验表明,该校准方法简化了鲁棒阈值的选择,显著提升了检测性能。

原文摘要 · Abstract (English)

Detecting occurrences of keywords with keyword spotting (KWS) systems requires thresholding continuous detection scores. Selecting appropriate thresholds is a non-trivial task, typically relying on optimizing performance on a validation dataset. However, such greedy threshold selection often leads to suboptimal performance on unseen data, particularly in varying or noisy acoustic environments or few-shot settings. In this work, we investigate detection threshold estimation for template-based open-set few-shot KWS using dynamic time warping on noisy speech data. To mitigate the performance degradation caused by suboptimal thresholds, we propose a score calibration approach that operates at the embedding level by quantizing learned representations and applying quantization error-based normalization prior to DTW-based scoring and thresholding. Experiments on KWS-DailyTalk with simulated high frequency radio channels show that the proposed calibration approach simplifies the selection of robust detection thresholds and significantly improves the resulting performance.

关键词检测少样本学习噪声鲁棒量化校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。