通过音素与语调协同学习,实现个性化关键词识别
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
- 双流编码器分别学习音素和说话人特有语调特征
- 融合模块动态结合两者信息,在多种语境下表现稳定
- 适合语音助手、个性语音控制等需要定制化识别的场景
现有关键词检测系统主要依赖音素级匹配来区分易混淆词汇,但忽略了用户特有的发音特征,如语调(语调、重音、节奏)。本文提出ProKWS框架,将细粒度音素学习与个性化语调建模相结合。设计双流编码器:一路径通过对比学习提取鲁棒的音素表示,另一路径提取说话人特异的语调模式。一个协同融合模块动态整合音素与语调信息,提升在不同声学环境下的适应性。实验表明,ProKWS在标准基准上达到与最先进模型相当的性能,并对带有语调和意图变化的个性化关键词表现出强鲁棒性。
原文摘要 · Abstract (English)
Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。