arXiv:2511.06296cs.SD2025-11

提出自监督混合训练框架,提升少样本语音关键词检测在混叠场景下的性能。

MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech

  • 在预训练阶段引入混合训练准则,自监督预测各信号的纯净声学单元。
  • 在谷歌语音命令数据集上,混叠与干净条件下均优于现有先进方法。
  • 适用于真实场景中多个关键词重叠出现的少样本语音识别任务。

少样本关键词检测旨在用极少量标注样本识别未见关键词。当前主流采用预训练+微调范式,但在混叠关键词检测(单个语句中同时存在多个重叠关键词)方面表现不佳,而该能力对实际应用至关重要。此前我们提出基于混合训练(MT)的预训练方法,有效应对混叠问题,但依赖全监督训练,无法利用大量无标签数据。为此,本文提出自监督学习框架 MT-HuBERT,将混合训练准则融入预训练过程。该方法通过上下文线索,自监督预测每个成分信号的纯净声学单元,而非混合语音的组合模式。在 Google Speech Commands (GSC v2) 数据集上的实验表明,所提方法在少样本关键词检测任务中,无论在混叠还是干净条件下,均持续优于多个前沿基线模型。

原文摘要 · Abstract (English)

Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. While effective in clean conditions, most existing approaches struggle with mixed keyword spotting--detecting multiple overlapping keywords within a single utterance--a capability essential for real-world applications. We have previously proposed a pre-training approach based on Mix-Training (MT) to tackle the mixed keyword detection problem and demonstrated its efficiency. However, this approach is fully supervised, unable to utilize vast unlabeled data. To this end, we propose Mix-Training HuBERT (MT-HuBERT), a self-supervised learning (SSL) pre-training framework that implements the MT criterion during pre-training. MT-HuBERT predicts, in a self-supervised manner, the clean acoustic units of each constituent signal from contextual cues, in contrast to predicting compositional patterns of mixed speech. Experiments conducted on the Google Speech Commands (GSC v2) corpus demonstrate that our proposed MT-HuBERT consistently outperforms several state-of-the-art baselines in few-shot KWS tasks under both mixed and clean conditions.

语音识别自监督学习关键词检测少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。