针对普通话口吃语音,构建检测与识别系统并验证效果。
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
- 基于AS-70数据集设计三类任务:口吃事件检测、口吃语音识别、创新方法研究。
- 最佳模型在口吃检测准确率与语音识别错误率上均有显著提升。
- 适合关注语音障碍技术、无障碍语音系统的研究人员与开发者。
口吃语音挑战赛致力于推动为口吃人群服务的语音技术发展,聚焦普通话口吃事件检测(SED)与口吃语音自动语音识别(ASR)。挑战赛包含三个赛道:(1)口吃事件检测,旨在开发口吃事件检测系统;(2)语音识别,专注于构建鲁棒的口吃语音识别系统;(3)研究赛道,鼓励利用所提供数据集提出创新方法。本研究采用开源普通话口吃数据集AS-70,已重新划分为新的训练集与测试集。本文介绍了该数据集,详述了挑战赛各赛道设置,并分析了顶尖系统的表现,突出显示检测准确率的提升与识别错误率的降低。研究结果强调了专用模型与增强策略在口吃语音技术开发中的潜力。
原文摘要 · Abstract (English)
The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recognition (ASR) in Mandarin. The challenge comprises three tracks: (1) SED, which aims to develop systems for detection of stuttering events; (2) ASR, which focuses on creating robust systems for recognizing stuttered speech; and (3) Research track for innovative approaches utilizing the provided dataset. We utilizes an open-source Mandarin stuttering dataset AS-70, which has been split into new training and test sets for the challenge. This paper presents the dataset, details the challenge tracks, and analyzes the performance of the top systems, highlighting improvements in detection accuracy and reductions in recognition error rates. Our findings underscore the potential of specialized models and augmentation strategies in developing stuttered speech technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。