arXiv:2605.22120eess.AScs.SD2026-05中稿 · TASLP被引 1

提出一种高效精准的自定义关键词唤醒框架,支持多模态注册与轻量持续学习。

Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation

论文配图:Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
图 1 · 摘自论文原文
  • 双阶段匹配:先定位候选片段,再用音素匹配精细验证,提升易混淆词区分度。
  • 在LibriPhrase Hard上达97.85% AUC、6.13% EER,优于现有方法。
  • 仅用18.7万参数即可实现轻量持续适配,适合设备端部署。

用户自定义关键词唤醒(KWS)对个性化语音交互至关重要,但现有方法存在三方面挑战:(1)易混淆词判别能力不足;(2)不同说话人发音差异导致性能不稳定;(3)需大量数据保障唤醒效果。本文提出DMA-KWS框架,通过三阶段优化解决上述问题:首先采用双阶段匹配流程——基于流式音素搜索的CTC解码定位候选段,再经QbyT音素匹配器进行细粒度验证,显著提升混淆词区分能力;其次引入多模态注册,融合用户语音与文本嵌入,增强注册用户识别精度;最后设计参数高效持续适应机制,利用合成与真实数据实现轻量级更新。大量实验表明,该方法在LibriPhrase Hard子集上达到97.85% AUC和6.13% EER,为当前最优水平。在说话人相关设置下,显著优于纯文本注册方式。此外,仅需187,000个可更新参数即可完成模型微调,兼顾性能提升与设备端部署可行性。

原文摘要 · Abstract (English)

User-defined keyword spotting (KWS) is crucial for personalized voice interaction, yet existing methods face several challenges: (1) insufficient discriminability among confusable words, (2) performance inconsistency across speakers with varying pronunciations, and (3) high data cost to ensure reliable wake-word performance. In this paper, we introduce DMA-KWS, an efficient and robust framework for user-defined keyword spotting. First, it adopts a dual-stage matching pipeline: CTC decoding with streaming phoneme search to locate candidate segments, followed by QbyT with a phoneme matcher for fine-grained verification, enabling it to better distinguish confusable words. Next, multi-modal enrollment fuses user-specific speech with text embeddings to further improve accuracy for registered users. Finally, a parameter-efficient continual adaptation mechanism performs lightweight updates using synthetic and real data. Extensive experiments demonstrate the superior performance of DMA-KWS. On the LibriPhrase Hard subset, it achieves 97.85% AUC and 6.13% EER, reaching state-of-the-art performance. In speaker-dependent settings, DMA-KWS consistently outperforms text-only enrollment, demonstrating significant performance gains. Moreover, the proposed parameter-efficient fine-tuning mechanism adapts DMA-KWS with only 187k updated parameters, further enhancing KWS performance while ensuring suitability for on-device deployment.

关键词唤醒多模态注册持续学习轻量化部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。