arXiv:2501.11631cs.SDcs.AI2025-01中稿 · ICASSP 2025被引 3

用预训练模型+噪声无关多任务学习,降低求助检测误报率

Noise-Agnostic Multitask Whisper Training for Reducing False Alarm Errors in Call-for-Help Detection

  • 基于预训练ASR模型,加入噪声分类头实现多任务学习
  • 在真实环境噪声下,误报率显著下降,检测性能提升
  • 无需重新训练即可适应新关键词,适合实际部署场景

关键词检测通常通过声学模型编码器中的关键词分类器实现,可应用于预定义或开放词汇的关键词识别。尽管该任务在各类应用中至关重要,并可扩展至紧急情况下的求助检测,但以往方法因引入新关键词或适应环境变化需重新训练,存在可扩展性局限。本文探索一种简单而有效的方法,利用现成的预训练自动语音识别(ASR)模型解决此问题,尤其适用于求助检测场景。此外,我们发现真实环境中麦克风噪声或不同环境导致求助检测系统误报率大幅上升。为此,提出一种新型噪声无关的多任务学习方法,在ASR编码器中集成噪声分类头。该方法提升了模型对噪声环境的鲁棒性,显著降低误报率并改善整体求助检测性能。尽管多任务学习带来一定复杂度,本方法计算高效,为真实场景下的求助检测提供了有前景的解决方案。

原文摘要 · Abstract (English)

Keyword spotting is often implemented by keyword classifier to the encoder in acoustic models, enabling the classification of predefined or open vocabulary keywords. Although keyword spotting is a crucial task in various applications and can be extended to call-for-help detection in emergencies, however, the previous method often suffers from scalability limitations due to retraining required to introduce new keywords or adapt to changing contexts. We explore a simple yet effective approach that leverages off-the-shelf pretrained ASR models to address these challenges, especially in call-for-help detection scenarios. Furthermore, we observed a substantial increase in false alarms when deploying call-for-help detection system in real-world scenarios due to noise introduced by microphones or different environments. To address this, we propose a novel noise-agnostic multitask learning approach that integrates a noise classification head into the ASR encoder. Our method enhances the model's robustness to noisy environments, leading to a significant reduction in false alarms and improved overall call-for-help performance. Despite the added complexity of multitask learning, our approach is computationally efficient and provides a promising solution for call-for-help detection in real-world scenarios.

语音识别求助检测多任务学习降噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。