arXiv:2505.19577eess.AScs.SD2025-05中稿 · TASLP被引 4

提出多头帧异步解码框架,显著提升关键词识别速度与准确率。

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding

  • 采用帧异步解码与专用声学模型,聚焦关键词检测。
  • 在多个数据集上达到当前最优性能,噪声环境下表现稳健。
  • 比传统方法快47%至63%,适合设备端部署。

关键词识别(KWS)对语音驱动应用至关重要,需兼顾准确性与效率。传统基于ASR的KWS方法如贪婪搜索和束搜索会遍历整个搜索空间,未显式优先考虑关键词检测,常导致性能不佳。本文提出一种面向流式处理的多头帧异步解码框架MFA-KWS,结合CTC-Transducer结构,采用关键词特定的音素同步解码,并以Token-and-Duration Transducer替代传统RNN-T,提升性能与效率。同时探索了单帧与一致性融合策略。大量实验表明,MFA-KWS在固定关键词与任意关键词数据集(如Snips、MobvoiHotwords、LibriKWS-20)上均达当前最优,且在噪声环境中表现鲁棒。其中,基于一致性的CDC-Last融合策略表现最佳。此外,相比帧同步基线,MFA-KWS在各数据集上提速47%至63%。实验验证其高效有效,适用于设备端部署。

原文摘要 · Abstract (English)

Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explore the entire search space without explicitly prioritizing keyword detection, often leading to suboptimal performance. In this paper, we propose an effective keyword-specific KWS framework by introducing a streaming-oriented CTC-Transducer-combined frame-asynchronous system with multi-head frame-asynchronous decoding (MFA-KWS). Specifically, MFA-KWS employs keyword-specific phone-synchronous decoding for CTC and replaces conventional RNN-T with Token-and-Duration Transducer to enhance both performance and efficiency. Furthermore, we explore various score fusion strategies, including single-frame-based and consistency-based methods. Extensive experiments demonstrate the superior performance of MFA-KWS, which achieves state-of-the-art results on both fixed keyword and arbitrary keywords datasets, such as Snips, MobvoiHotwords, and LibriKWS-20, while exhibiting strong robustness in noisy environments. Among fusion strategies, the consistency-based CDC-Last method delivers the best performance. Additionally, MFA-KWS achieves a 47% to 63% speed-up over the frame-synchronous baselines across various datasets. Extensive experimental results confirm that MFA-KWS is an effective and efficient KWS framework, making it well-suited for on-device deployment.

关键词识别语音处理流式解码模型加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。