arXiv:2410.05647cs.SDeess.AS2024-10中稿 · SLT 2024被引 1

通过细粒度对比学习提升普通话口吃事件检测准确率

FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection

  • 基于帧级概率建模,挖掘易混淆语音帧
  • 在中文数据上F1得分提升超5.0%
  • 适合语音分析与口吃辅助诊断研究者

本文介绍了T031团队在SLT2024口吃语音挑战赛中的方法。普通话口吃事件检测(MSED)旨在识别普通话语音中的口吃片段。为提高检测精度,我们提出一种细致的声学分析方法,捕捉以往口吃事件检测(SED)技术忽略的细微特征。为此,我们引入细粒度对比学习(FGCL)框架用于MSED:首先建模帧级口吃概率,设计挖掘算法识别易样本与混淆样本;随后提出口吃对比损失,增强口吃与流畅语音帧之间的区分能力,提升特征嵌入的判别性能。在英、中语种数据集上的大量实验表明,该方法有效,中文数据上F1分数提升超过5.0%。

原文摘要 · Abstract (English)

This paper presents the T031 team's approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data.

口吃检测对比学习语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。