arXiv:2509.19879eess.AS2025-09被引 2

用弱监督方法提取可解释的发音特征,助力病理语音分析

Weakly Supervised Phonological Features for Pathological Speech Analysis

  • 通过语音识别模型引入可解释的发音特征瓶颈层
  • 病理语音分类准确率达75%,流畅度预测RMSE为8.43
  • 特征不依赖文本且可解释,适合临床医生使用

语音的副语言特性在评估和选择言语障碍患者的治疗方案中至关重要。然而,由于缺乏标注了副语言特性的语音数据集,尤其是在帧级标注方面,自动建模这些特性十分困难。本文提出一种弱监督训练方法,利用已知的音素声学特性,通过在语音识别模型中加入可解释的帧级发音特征瓶颈层进行训练。随后,我们构建了用于语音可懂度预测和病理语音分类的模型,评估所提取发音特征的有效性。实验结果显示,使用该方法的模型在两项任务上的表现均达到当前先进水平:分类准确率为75%,语音可懂度预测的RMSE为8.43。与现有方法相比,本方法提取的发音特征具有无需依赖文本、高度可解释的优点,可为语音治疗师提供潜在有价值的临床洞察。

原文摘要 · Abstract (English)

Paralinguistic properties of speech are essential in analyzing and choosing optimal treatment options for patients with speech disorders. However, automatic modeling of these characteristics is difficult due to the lack of labeled speech datasets describing paralinguistic properties, especially at the frame-level. In this paper, we propose a weakly supervised training method which exploits the known acoustic properties of phonemes by training an ASR model with an interpretable frame-level phonological feature bottleneck layer. Subsequently, we assess the viability of these phonological features in speech pathology analysis by developing corresponding models for intelligibility prediction and speech pathology classification. Models using our proposed phonological features perform similar to other state-of-the-art acoustic features on both tasks with a classification accuracy of 75% and a 8.43 RMSE on speech intelligibility prediction. In contrast to others, our phonological features are text-independent and highly interpretable, providing potentially useful insights for speech therapists.

语音分析弱监督可解释性病理语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。