用关键帧融合提升自定义关键词识别准确率
KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

- 通过CTC识别高置信度音素帧,实现多模态精准对齐
- 在难例集上达97.65% AUC和7.75% EER,性能显著领先
- 适合需要高精度个性化语音交互的场景
用户自定义关键词检测(KWS)通过识别用户指定关键词,实现个性化语音交互。该任务的关键挑战在于区分语音相似的关键词。为此,我们提出KFC-KWS,一种基于连接时序分类(CTC)引导的关键帧选择的多模态框架。具体地,利用CTC产生的尖峰后验分布识别高置信度音素帧,实现音频、音素与文本模态间的精确对齐。这些关键帧通过交叉注意力与完整语句表示融合,同时捕捉局部判别特征与全局上下文信息。在LibriPhrase数据集上,KFC-KWS取得最佳平衡性能(98.73% AUC),在具有挑战性的难例子集上显著优于先进基线模型(97.65% AUC 和 7.75% EER),验证了其在高度混淆关键词间区分能力的有效性。
原文摘要 · Abstract (English)
User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from phonetically confusable alternatives. To address this challenge, we propose KFC-KWS, a multimodal framework that leverages connectionist temporal classification (CTC)-guided keyframe selection. Specifically, we exploit the peaky posterior distributions of CTC to identify high-confidence phoneme frames, enabling precise alignment across audio, phoneme, and text modalities. These keyframes are then fused with full-utterance representations through cross-attention to capture both local discriminative cues and global contextual information. On LibriPhrase, KFC-KWS achieves the best-balanced performance (98.73% AUC) and substantially outperforms advanced baselines on the challenging hard subset (97.65% AUC and 7.75% EER), demonstrating its effectiveness in discriminating between highly confusable keywords.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。