arXiv:2603.18023eess.AScs.AI2026-03

轻量级多任务模型实现个性化关键词唤醒,兼顾隐私与效率

PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting

  • 用轻量网络同时做关键词识别和说话人验证
  • 新损失函数避免类别间竞争,提升识别准确率
  • 参数少、算力低,适合嵌入式设备部署

随着物联网(IoT)、自动语音识别(ASR)、说话人验证(SV)和文本转语音(TTS)等技术的发展,智能语音助手的使用日益普及,对隐私性和个性化的需求也显著上升。本文提出一种用于个性化可定制开放词汇关键词唤醒(PCOV-KWS)的多任务学习框架。该框架采用轻量级网络,同时执行关键词唤醒(KWS)与说话人验证(SV),以满足个性化唤醒需求。引入不同于Softmax的训练准则,将多分类问题转化为多个二分类问题,消除了类别间的竞争关系;同时在训练中采用多任务损失加权优化策略。我们在多个数据集上评估了PCOV-KWS系统,结果表明其优于基线方法,且所需参数更少、计算资源更低。

原文摘要 · Abstract (English)

As advancements in technologies like Internet of Things (IoT), Automatic Speech Recognition (ASR), Speaker Verification (SV), and Text-to-Speech (TTS) lead to increased usage of intelligent voice assistants, the demand for privacy and personalization has escalated. In this paper, we introduce a multi-task learning framework for personalized, customizable open-vocabulary Keyword Spotting (PCOV-KWS). This framework employs a lightweight network to simultaneously perform Keyword Spotting (KWS) and SV to address personalized KWS requirements. We have integrated a training criterion distinct from softmax-based loss, transforming multi-class classification into multiple binary classifications, which eliminates inter-category competition, while an optimization strategy for multi-task loss weighting is employed during training. We evaluated our PCOV-KWS system in multiple datasets, demonstrating that it outperforms the baselines in evaluation results, while also requiring fewer parameters and lower computational resources.

关键词唤醒多任务学习个性化语音轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。