arXiv:2502.00295eess.AScs.SD2025-02

用耳机内麦+课程学习,让耳语关键词识别更抗噪

Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning

  • 用耳机内置麦克风采集耳语语音,结合多麦克风处理
  • 课程学习策略逐步提升耳语数据比例,使噪声下F1提升15%
  • 适合开发静音场景下语音控制的智能耳机产品

现代耳机功能日益丰富,传统物理按键难以应对复杂操作。语音指令中的关键词识别成为替代方案,但现有方法仅支持常规语音,而常规语音在安静或公共场合不适用。本文研究设备端耳语关键词识别问题,利用降噪耳机内的麦克风作为额外语音输入源,并设计课程学习策略,在训练中逐步增加耳语关键词的比例。实验表明,多麦克风处理与课程学习结合可使噪声环境下耳语关键词识别的F1分数提升最高达15%。

原文摘要 · Abstract (English)

The expanding feature set of modern headphones puts a challenge on the design of their control interface. Users may want to separately control each feature or quickly switch between modes that activate different features. Traditional approach of physical buttons may no longer be feasible when the feature set is large. Keyword spotting with voice commands is a promising solution to the issue. Most existing methods of keyword spotting only support commands spoken in a regular voice. However, regular voice may not be desirable in quiet places or public settings. In this paper, we investigate the problem of on-device keyword spotting in whisper voice and explore approaches to improve noise robustness. We leverage the inner microphone on noise-cancellation headphones as an additional source of voice input. We also design a curriculum learning strategy that gradually increases the proportion of whisper keywords during training. We demonstrate through experiments that the combination of multi-microphone processing and curriculum learning could improve F1 score of whisper keyword spotting by up to 15% in noisy conditions.

关键词识别耳语语音降噪耳机课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。