arXiv:2508.21248eess.AScs.AI2025-08中稿 · ed被引 3

用自监督模型特征实现儿童语音零样本关键词识别

Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models

  • 从SSL模型多层提取特征,构建零样本关键词识别系统
  • 在30个关键词上达到0.691的ATWV得分,误报率仅0.0164
  • 适用于不同年龄儿童,对噪声环境也有良好鲁棒性

针对儿童语音在声学与语言特征上与成人差异显著的问题,本文提出一种零样本关键词识别(KWS)方法,利用Wav2Vec2、HuBERT和Data2Vec等先进自监督学习(SSL)模型的分层特征。使用WSJCAM0数据集训练,以PFSTAR儿童语音数据集测试,验证了零样本能力。结果表明,该方法在所有关键词集合上均达当前最优表现;其中Wav2Vec2第22层效果最佳,30个关键词下ATWV为0.691,MTWV为0.7003,误报率0.0164,漏检率0.0547。年龄分组评估显示系统对不同年龄段儿童均有效。在噪声环境下,基于最优层的实验仍显著优于传统MFCC基线。进一步在CMU数据集上复现,验证了框架的泛化能力。结果表明SSL特征可有效提升儿童语音零样本KWS性能。

原文摘要 · Abstract (English)

Numerous methods have been proposed to enhance Keyword Spotting (KWS) in adult speech, but children's speech presents unique challenges for KWS systems due to its distinct acoustic and linguistic characteristics. This paper introduces a zero-shot KWS approach that leverages state-of-the-art self-supervised learning (SSL) models, including Wav2Vec2, HuBERT and Data2Vec. Features are extracted layer-wise from these SSL models and used to train a Kaldi-based DNN KWS system. The WSJCAM0 adult speech dataset was used for training, while the PFSTAR children's speech dataset was used for testing, demonstrating the zero-shot capability of our method. Our approach achieved state-of-the-art results across all keyword sets for children's speech. Notably, the Wav2Vec2 model, particularly layer 22, performed the best, delivering an ATWV score of 0.691, a MTWV score of 0.7003 and probability of false alarm and probability of miss of 0.0164 and 0.0547 respectively, for a set of 30 keywords. Furthermore, age-specific performance evaluation confirmed the system's effectiveness across different age groups of children. To assess the system's robustness against noise, additional experiments were conducted using the best-performing layer of the best-performing Wav2Vec2 model. The results demonstrated a significant improvement over traditional MFCC-based baseline, emphasizing the potential of SSL embeddings even in noisy conditions. To further generalize the KWS framework, the experiments were repeated for an additional CMU dataset. Overall the results highlight the significant contribution of SSL features in enhancing Zero-Shot KWS performance for children's speech, effectively addressing the challenges associated with the distinct characteristics of child speakers.

语音识别儿童语音零样本自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。