用语音识别模型提取发音与语调特征,提升构音障碍严重程度自动评估准确率。
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
- 用微调的ASR模型转录失语语音,提取词段边界信息作为新特征
- 在临床数据集上达到83.72%的平衡准确率,优于现有方法
- 特征可解释性强,适合临床辅助诊断场景
由于当前临床评估具有主观性,自动评估构音障碍语音严重程度的需求日益迫切。深度神经网络(DNN)模型性能优于传统机器学习(ML)模型,但缺乏可解释性;而传统机器学习模型虽可在特征层面提供可解释结果,但性能相对较低。现有机器学习方法从原始波形中提取多种特征以预测严重程度,但未涵盖临床评估中所有构音障碍特征。为填补此空白,本文提出一种最小化信息损失的特征提取方法,引入语音识别(ASR)转录作为新型特征来源。通过微调ASR模型以适应构音障碍语音,利用该模型对失语语音进行转录,并提取词段边界信息,从而捕捉更精细的发音特征与更广泛的语调特征。实验表明,该方法在严重程度预测任务中表现优于现有特征,达到83.72%的平衡准确率。
原文摘要 · Abstract (English)
Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable results at a feature level, but their performance is comparatively lower. Current ML models extract various features from raw waveforms to predict severity. However, existing methods do not encompass all dysarthric features used in clinical evaluation. To address this gap, we propose a feature extraction method that minimizes information loss. We introduce an ASR transcription as a novel feature extraction source. We finetune the ASR model for dysarthric speech, then use this model to transcribe dysarthric speech and extract word segment boundary information. It enables capturing finer pronunciation and broader prosodic features. These features demonstrated an improved severity prediction performance to existing features: balanced accuracy of 83.72%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。