无需训练数据,通过语音表征中的发音特征空间变化评估失语症严重程度。
Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations
- 利用健康人语音数据提取发音特征方向,冻结模型不训练。
- 12维发音特征在10个语料库中与临床严重度显著相关(ρ=-0.47至-0.55)。
- 适用于29种语言,适合无标注数据的跨语言失语评估研究。
失语症严重程度评估通常依赖受过训练的临床医生或带标签病理语音的监督模型,限制了其在多语言和临床场景中的可扩展性。本文提出一种无需训练的方法,通过测量冻结的HuBERT表示中发音特征子空间的退化程度来量化失语症严重程度。特征方向仅从健康对照语音中通过预训练的强制对齐器估计。对每位说话者,使用蒙特利尔强制对齐器提取音素级嵌入,计算沿发音对比方向(鼻音性、浊音性、强音性、响度、方式及4个元音特征)的d-prime得分,构建12维发音特征谱。在涵盖890名说话者、5种语言(英语、西班牙语、荷兰语、汉语、法语)和3种主要病因(帕金森病、脑瘫、肌萎缩侧索硬化症)的10个语料库上评估,所有5个辅音d-prime特征均与临床严重度显著相关(随机效应元分析ρ = -0.50至-0.56,p < 2e-4;合并斯皮尔曼ρ = -0.47至-0.55,置信区间不包含零)。该效应在单个语料库内复现,经FDR校正后仍显著,且对去除单个语料库和对齐质量控制保持稳健。鼻音性d-prime在7个分级语料库中的6个呈现从健康到重度的单调下降。曼-惠特尼U检验表明所有12个特征均可有效区分健康者与重度失语者(p < 0.001)。该方法无需失语症训练数据,适用于已有MFA声学模型的任意语言(目前支持29种语言)。我们公开完整流程与6种语言的音素特征配置。
原文摘要 · Abstract (English)
Dysarthric speech severity assessment typically requires trained clinicians or supervised models built from labelled pathological speech, limiting scalability across languages and clinical settings. We present a training-free method that quantifies dysarthria severity by measuring degradation in phonological feature subspaces within frozen HuBERT representations. No supervised severity model is trained; feature directions are estimated from healthy control speech using a pretrained forced aligner. For each speaker, we extract phone-level embeddings via Montreal Forced Aligner, compute d-prime scores along phonological contrast directions (nasality, voicing, stridency, sonorance, manner, and four vowel features) derived exclusively from healthy controls, and construct a 12-dimensional phonological profile.Evaluating 890 speakers across 10 corpora, 5 languages (English, Spanish, Dutch, Mandarin, French), and 3 primary aetiologies (Parkinson's disease, cerebral palsy, ALS), we find that all five consonant d-prime features correlate significantly with clinical severity (random-effects meta-analysis rho = -0.50 to -0.56, p < 2e-4; pooled Spearman rho = -0.47 to -0.55 with bootstrap 95% CIs not crossing zero). The effect replicates within individual corpora, survives FDR correction, and remains robust to leave-one-corpus-out removal and alignment quality controls. Nasality d-prime decreases monotonically from control to severe in 6 of 7 severity-graded corpora. Mann-Whitney U tests confirm that all 12 features distinguish controls from severely dysarthric speakers (p < 0.001).The method requires no dysarthric training data and applies to any language with an existing MFA acoustic model (currently 29 languages). We release the full pipeline and phone feature configurations for six languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。