arXiv:2409.10076cs.SDcs.HC2024-09中稿 · SLT 2024被引 1

针对口吃者语音唤醒难题,提出端到端双滤波系统,显著降低误唤醒率。

Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge

  • 采用多任务微调的双分支data2vec2模型,统一建模语音识别与唤醒词检测
  • 在测试集上实现0.321%误唤醒率、0.5%漏唤醒率,总分0.821%
  • 适合残障语音交互、低资源口吃语音处理场景

语音已成为各类应用中广泛采用的用户界面。然而,对于患有构音障碍的人群,其语音固有的不稳定性带来了显著挑战。本文针对SLT 2024低资源构音障碍语音唤醒词检测挑战赛,提出一种基于预训练的端到端双滤波构音障碍语音唤醒系统(PD-DWS)。该系统从音频建模与双滤波策略两方面提升性能:在音频建模方面,提出基于预训练data2vec2(d2v2)的双分支2branch-d2v2模型,通过统一的多任务微调范式,同时建模自动语音识别(ASR)与语音唤醒词检测(WWS)任务;此外,引入双滤波策略,在保持相同漏唤醒率(FRR)的前提下,有效降低误唤醒率(FAR)。实验结果表明,所提PD-DWS系统在测试集B上达到FAR为0.00321、FRR为0.005,总分为0.00821,获得挑战赛第一名。

原文摘要 · Abstract (English)

Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses significant challenges. This paper presents an end-to-end Pretrain-based Dual-filter Dysarthria Wake-up word Spotting (PD-DWS) system for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge. Specifically, our system improves performance from two key perspectives: audio modeling and dual-filter strategy. For audio modeling, we propose an innovative 2branch-d2v2 model based on the pre-trained data2vec2 (d2v2), which can simultaneously model automatic speech recognition (ASR) and wake-up word spotting (WWS) tasks through a unified multi-task finetuning paradigm. Additionally, a dual-filter strategy is introduced to reduce the false accept rate (FAR) while maintaining the same false reject rate (FRR). Experimental results demonstrate that our PD-DWS system achieves an FAR of 0.00321 and an FRR of 0.005, with a total score of 0.00821 on the test-B eval set, securing first place in the challenge.

语音唤醒口吃语音多任务学习低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。