arXiv:2409.10704eess.AScs.AI2024-09中稿 · IEEE SLT 2024被引 10

用自监督模型实现单词级口吃语音检测,提升筛查效率

Self-supervised Speech Models for Word-Level Stuttered Speech Detection

  • 基于自监督语音模型构建单词级口吃检测方法
  • 在自建标注数据集上性能超越已有方法
  • 适合语音病理学研究与自动化筛查应用

口吃临床诊断需由受过专业训练的言语语言病理学家完成,但该过程耗时且专业人才稀缺。全球约8000万人口吃,而具备相关经验的专家比例极低。开发机器学习模型自动检测口吃语音,可实现普遍化、自动化筛查,帮助病理学家识别高风险患者。以往研究多聚焦语句级检测,难以满足临床中以单词为单位标注的需求。本研究构建了首个带单词级标注的口吃语音数据集,并提出一种基于自监督语音模型的单词级口吃检测方法。实验表明,该模型在单词级检测任务上显著优于现有方法。此外,通过系统消融分析,揭示了自监督模型适配口吃检测的关键因素。

原文摘要 · Abstract (English)

Clinical diagnosis of stuttering requires an assessment by a licensed speech-language pathologist. However, this process is time-consuming and requires clinicians with training and experience in stuttering and fluency disorders. Unfortunately, only a small percentage of speech-language pathologists report being comfortable working with individuals who stutter, which is inadequate to accommodate for the 80 million individuals who stutter worldwide. Developing machine learning models for detecting stuttered speech would enable universal and automated screening for stuttering, enabling speech pathologists to identify and follow up with patients who are most likely to be diagnosed with a stuttering speech disorder. Previous research in this area has predominantly focused on utterance-level detection, which is not sufficient for clinical settings where word-level annotation of stuttering is the norm. In this study, we curated a stuttered speech dataset with word-level annotations and introduced a word-level stuttering speech detection model leveraging self-supervised speech models. Our evaluation demonstrates that our model surpasses previous approaches in word-level stuttering speech detection. Additionally, we conducted an extensive ablation analysis of our method, providing insight into the most important aspects of adapting self-supervised speech models for stuttered speech detection.

语音检测自监督口吃识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。