arXiv:2510.10740cs.SD2025-10被引 5

通过双阶段框架与数据扩增,实现高鲁棒性用户自定义关键词检测。

Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting

  • 两阶段设计:先定位候选片段,再逐音素与整句验证。
  • 在LibriPhrase硬集上达6.13%错误率、97.85%准确率。
  • 零样本性能媲美全量训练,适合语音助手等实际应用。

本文提出DS-KWS,一种针对鲁棒性用户自定义关键词探测的两阶段框架。该方法结合基于CTC的方法与流式音素搜索模块以定位候选片段,随后采用基于QbyT的方法与音素匹配模块,在音素和语句层面进行验证。为进一步提升性能,引入双数据扩增策略:(1) 将ASR语料库从460小时扩展至1,460小时,强化声学模型;(2) 利用超过15.5万个锚类别训练音素匹配器,显著增强易混淆词的区分能力。在LibriPhrase数据集上,DS-KWS显著优于现有方法,于Hard子集达到6.13% EER与97.85% AUC。在Hey-Snips上,其零样本性能可媲美全量训练模型,实现每小时仅1次误报下的99.13%召回率。

原文摘要 · Abstract (English)

In this paper, we propose DS-KWS, a two-stage framework for robust user-defined keyword spotting. It combines a CTC-based method with a streaming phoneme search module to locate candidate segments, followed by a QbyT-based method with a phoneme matcher module for verification at both the phoneme and utterance levels. To further improve performance, we introduce a dual data scaling strategy: (1) expanding the ASR corpus from 460 to 1,460 hours to strengthen the acoustic model; and (2) leveraging over 155k anchor classes to train the phoneme matcher, significantly enhancing the distinction of confusable words. Experiments on LibriPhrase show that DS-KWS significantly outperforms existing methods, achieving 6.13\% EER and 97.85\% AUC on the Hard subset. On Hey-Snips, it achieves zero-shot performance comparable to full-shot trained models, reaching 99.13\% recall at one false alarm per hour.

关键词检测语音识别零样本音素匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。