arXiv:2606.10365cs.SD2026-06中稿 · Interspeech 2026

用关键帧融合提升自定义关键词识别准确率

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

论文配图:KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting
图 1 · 摘自论文原文
  • 通过CTC识别高置信度音素帧,实现多模态精准对齐
  • 在难例集上达97.65% AUC和7.75% EER,性能显著领先
  • 适合需要高精度个性化语音交互的场景

用户自定义关键词检测(KWS)通过识别用户指定关键词,实现个性化语音交互。该任务的关键挑战在于区分语音相似的关键词。为此,我们提出KFC-KWS,一种基于连接时序分类(CTC)引导的关键帧选择的多模态框架。具体地,利用CTC产生的尖峰后验分布识别高置信度音素帧,实现音频、音素与文本模态间的精确对齐。这些关键帧通过交叉注意力与完整语句表示融合,同时捕捉局部判别特征与全局上下文信息。在LibriPhrase数据集上,KFC-KWS取得最佳平衡性能(98.73% AUC),在具有挑战性的难例子集上显著优于先进基线模型(97.65% AUC 和 7.75% EER),验证了其在高度混淆关键词间区分能力的有效性。

原文摘要 · Abstract (English)

User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from phonetically confusable alternatives. To address this challenge, we propose KFC-KWS, a multimodal framework that leverages connectionist temporal classification (CTC)-guided keyframe selection. Specifically, we exploit the peaky posterior distributions of CTC to identify high-confidence phoneme frames, enabling precise alignment across audio, phoneme, and text modalities. These keyframes are then fused with full-utterance representations through cross-attention to capture both local discriminative cues and global contextual information. On LibriPhrase, KFC-KWS achieves the best-balanced performance (98.73% AUC) and substantially outperforms advanced baselines on the challenging hard subset (97.65% AUC and 7.75% EER), demonstrating its effectiveness in discriminating between highly confusable keywords.

关键词检测多模态CTC语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。