arXiv:2604.03689eess.AS2026-04中稿 · ICASSP 2026

零样本关键词识别新框架,精准抑制误报,轻量高效适合设备端部署。

MALEFA: Multi-grAnularity Learning and Effective False Alarm Suppression for Zero-shot Keyword Spotting

  • 通过跨注意力联合学习话语与音素级对齐,提升语义匹配精度。
  • 在AMI数据集上误报率低至0.007%,准确率达90%。
  • 模型轻量,适配资源受限设备,支持实时运行。

用户自定义关键词检测(KWS)无需领域特定标注训练数据,对构建可适应、个性化的语音交互系统至关重要。然而,此类系统仍面临计算资源有限和标注数据匮乏的挑战。现有方法难以区分声学相似关键词,常导致真实场景部署中误报率高企。为此,我们提出MALEFA,一种新颖的轻量化零样本KWS框架,通过交叉注意力机制联合学习话语级与音素级对齐,并引入多粒度对比学习目标。在四个公开基准数据集上的评估显示,MALEFA在准确性达90%的同时,将AMI数据集上的误报率显著降低至0.007%。除优异性能外,MALEFA还展现出高计算效率,可直接支持资源受限设备上的实时部署。

原文摘要 · Abstract (English)

User-defined keyword spotting (KWS) without resorting to domain-specific pre-labeled training data is of fundamental importance in building adaptable and personalized voice interfaces. However, such systems are still faced with arduous challenges, including constrained computational resources and limited annotated training data. Existing methods also struggle to distinguish acoustically similar keywords, often leading to a pesky false alarm rate (FAR) in real-world deployments. To mitigate these limitations, we put forward MALEFA, a novel lightweight zero-shot KWS framework that jointly learns utterance- and phoneme-level alignments via cross-attention and a multi-granularity contrastive learning objective. Evaluations on four public benchmark datasets show that MALEFA achieves a high accuracy of 90%, significantly reducing FAR to 0.007% on the AMI dataset. Beyond its strong performance, MALEFA demonstrates high computational efficiency and can readily support real-time deployment on resource-constrained devices.

关键词识别零样本轻量化误报抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。