arXiv:2506.00454eess.AScs.HC2025-06中稿 · Interspeech 2025被引 2

用自动语音识别技术分析口吃言语清晰度,实现逐时错误定位与分类。

Towards Temporally Explainable Dysarthric Speech Clarity Assessment

  • 构建三阶段框架:整体清晰度评分、错误定位、类型分类
  • 基于6名说话者读两段文字的数据集,实现可解释的发音错误分析
  • 适合语言治疗师与语音康复研究者参考,推动个性化反馈系统

口吃是一种运动性言语障碍,影响言语可理解性,需针对性干预以促进有效沟通。本文通过收集6位说话者朗读两段文字的口吃语音数据集,并由言语治疗师标注时间标记和发音错误描述,开展自动化发音错误反馈研究。设计了一个三阶段可解释的发音错误评估框架:(1) 整体清晰度评分,(2) 发音错误定位,(3) 发音错误类型分类。系统评估了预训练自动语音识别(ASR)模型在各阶段的表现,验证其在口吃言语评估中的有效性(代码地址:https://github.com/augmented-human-lab/interspeech25_speechtherapy,补充网页:https://apps.ahlab.org/interspeech25_speechtherapy/)。研究结果为自动化可操作的发音评估反馈提供了临床相关洞见,有助于患者独立练习并提升治疗师干预效率。

原文摘要 · Abstract (English)

Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric speech dataset from six speakers reading two passages, annotated by a speech therapist with temporal markers and mispronunciation descriptions. We design a three-stage framework for explainable mispronunciation evaluation: (1) overall clarity scoring, (2) mispronunciation localization, and (3) mispronunciation type classification. We systematically analyze pretrained Automatic Speech Recognition (ASR) models in each stage, assessing their effectiveness in dysarthric speech evaluation (Code available at: https://github.com/augmented-human-lab/interspeech25_speechtherapy, Supplementary webpage: https://apps.ahlab.org/interspeech25_speechtherapy/). Our findings offer clinically relevant insights for automating actionable feedback for pronunciation assessment, which could enable independent practice for patients and help therapists deliver more effective interventions.

语音评估口吃可解释性ASR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。