arXiv:2606.20106eess.AScs.SD2026-06中稿 · Interspeech 2026

让语音唤醒支持用户自定义关键词和说话人验证,零样本部署

Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

论文配图:Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification
图 1 · 摘自论文原文
  • 用音素监督音频编码+预训练说话人编码,实现双零样本识别
  • 在多个数据集上误拒率降低60%,参数仅155万适合边缘设备
  • 无需重新训练即可切换普通唤醒或严格说话人认证模式

用户自定义关键词检测(UD-KWS)实现了基于文本的零样本唤醒词检测,但现有系统学习的是与说话人无关的表征,无法拒绝正确关键词但非本人说出的情况。本文提出ZP-KWS,一个轻量级框架,结合音素监督音频编码器与GE2E预训练的紧凑说话人编码器(约0.9M参数)。推理时采用乘法式晚期融合,使各分支拥有独立否决权,支持从常规检测到严格说话人门控激活的多种模式,无需重新训练。在LibriPhrase、Google Speech Commands和Qualcomm数据集上,ZP-KWS在1%误报率下目标仅误拒率相对基线最高降低60%,同时保持优异的关键词检测性能,整体参数量控制在155万以内,适用于边缘部署。

原文摘要 · Abstract (English)

User-defined keyword spotting (UD-KWS) enables zero-shot wake-word detection from text, but existing systems learn speaker-invariant representations that cannot reject impostors uttering the correct keyword. We address this dual zero-shot setting -- unseen keywords and unseen speakers -- with ZP-KWS, a lightweight framework combining a phoneme-supervised audio encoder with a GE2E-pretrained compact speaker encoder (about 0.9M parameters). Multiplicative late fusion at inference grants each branch independent veto power, supporting modes from conventional detection to strict speaker-gated activation without retraining. On LibriPhrase, Google Speech Commands, and Qualcomm datasets, ZP-KWS reduces target-only FRR at 1% FAR by up to 60% relative to the strongest baseline while maintaining competitive keyword detection, all within a 1.55M parameter budget for edge deployment.

语音唤醒说话人验证轻量化模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。