arXiv:2412.20805eess.AScs.SD2024-12被引 18

通过音素级对比学习,提升自定义关键词识别的准确率与灵活性。

Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment

  • 在音素层级进行对比学习,增强语音与文本特征对齐精度。
  • 在LibriPhrase数据集上实现当前最优性能,误报率显著降低。
  • 支持多模态录入,适合需要灵活定制关键词的应用场景。

用户自定义关键词检测(KWS)通过允许用户自定义关键词来提升体验。然而,在开放词汇场景下,现有方法普遍面临易混淆词导致的高误报率问题,且仅支持音频或文本单一模式录入。为此,本文首次探索模型对易混淆词的鲁棒性,提出音素级对比学习(PLCL),在音素层级细化并对齐查询与源特征表示。该方法通过细粒度正负样本对比,提升模型区分能力,并可统一优化音频-文本与音频-音频匹配,适配多种录入方式。此外,构建无上下文音素记忆库以生成混淆负样本用于数据增强,并设计三类判别器专门区分困难负样本。整体上,建立了一个鲁棒且灵活的KWS系统,支持统一框架下的多模态录入。在LibriPhrase数据集上的实验验证了其先进性能。

原文摘要 · Abstract (English)

User-defined keyword spotting (KWS) enhances the user experience by allowing individuals to customize keywords. However, in open-vocabulary scenarios, most existing methods commonly suffer from high false alarm rates with confusable words and are limited to either audio-only or text-only enrollment. Therefore, in this paper, we first explore the model's robustness against confusable words. Specifically, we propose Phoneme-Level Contrastive Learning (PLCL), which refines and aligns query and source feature representations at the phoneme level. This method enhances the model's disambiguation capability through fine-grained positive and negative comparisons for more accurate alignment, and it is generalizable to jointly optimize both audio-text and audio-audio matching, adapting to various enrollment modes. Furthermore, we maintain a context-agnostic phoneme memory bank to construct confusable negatives for data augmentation. Based on this, a third-category discriminator is specifically designed to distinguish hard negatives. Overall, we develop a robust and flexible KWS system, supporting different modality enrollment methods within a unified framework. Verified on the LibriPhrase dataset, the proposed approach achieves state-of-the-art performance.

关键词检测音素级学习多模态对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。