arXiv:2608.28640cs.CLcs.AI2026-09

用提示词提升关键词唤醒准确率,尤其在嘈杂环境表现更优

PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

论文配图:PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework
图 1 · 摘自论文原文
  • 设计提示词预测网络,从关键词生成语义嵌入
  • 通过跨注意力机制融合提示与声学特征,唤醒率提升超10%
  • 对噪声和发音差异鲁棒,适合真实复杂场景应用

本文提出 PromptKWS,一种新型提示词引导的开放词汇关键词唤醒框架,旨在提升开放词汇 KWS 系统的准确性。具体而言,我们引入提示词短语预测网络(PPN),采用编码器-解码器架构,有效提取关键词提示的嵌入表示。利用 PPN 编码器对关键词提示进行编码,并通过提示-声学多头交叉注意力(MHCA)将提示嵌入注入提示引导的 KWS 编码器中。实验表明,与基线系统相比,PromptKWS 在唤醒率上提升超过 10%。值得注意的是,该方法在应对噪声和发音变化等复杂现实环境时表现出色,相较于纯声学模型,在测试集上平均准确率提升超过 15%。

原文摘要 · Abstract (English)

In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder architecture designed to effectively extract keyword prompts embeddings. we employ the PPN encoder to encode the keyword prompts and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing a Prompt-acoustic Multi-head Cross-attention (MHCA). Experiments show that PromptKWS improves the wakeup rate by over 10% compared to baseline system. Notably, another strength of PromptKWS is its ability to effectively leverage keyword prompts for adapting to complex real-world environments involving noise and pronunciation variations. In comparison to purely acoustic models, which often struggle in such situations, PromptKWS demonstrates remarkable performance, with an average accuracy improvement of over 15% in test sets.

关键词唤醒提示学习语音识别鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。