arXiv:2604.21706cs.CL2026-04

无需训练,通过语音表征分析可识别不同病因导致的发音障碍模式。

Phonological Subspace Collapse Is Aetiology-Specific and Cross-Lingually Stable: Evidence from 3,374 Speakers

论文配图:Phonological Subspace Collapse Is Aetiology-Specific and Cross-Lingually Stable: Evidence from 3,374 Speakers
图 1 · 摘自论文原文
  • 基于冻结自监督语音表征的音位子空间分离度,实现无训练评估发音障碍严重程度。
  • 13个音位特征中有10个显示显著差异,帕金森病与运动执行障碍组可区分(效应量0.83)。
  • 跨语言特征形状稳定,6种模型均表现一致,适合多语言、多病因研究使用。

此前我们提出一种无需训练的方法,通过冻结自监督语音表示中音位特征子空间的d-prime可分性,评估构音障碍严重程度,在5种语言共890名受试者上验证了HuBERT-base的有效性。本文将分析扩展至25个数据集、3,374名受试者,覆盖12种语言和5种病因(帕金森病、脑瘫、肌萎缩侧索硬化症、唐氏综合征、中风),以及健康对照组,使用6种自监督学习骨干模型。结果表明:第一,病因特异性退化模式在群体层面可区分,13个特征中有10个产生大效应量(epsilon-squared > 0.14,Holm校正后p < 0.001),帕金森病与运动执行障碍组间差异达Cohen's d = 0.83;个体分类性能有限(宏平均F1为22.6%)。第二,特征轮廓在跨语言间保持高稳定性,每种病因下5维辅音d-prime轮廓的余弦相似度超过0.95;但绝对d-prime值不具跨语言可比性,方法支持语言无关的退化模式表型,但需在单语库内校准绝对严重程度。第三,该方法对模型架构不敏感:所有6个骨干模型均呈现单调严重程度梯度,模型间一致性高于rho = 0.77。固定词元数量下的d-prime估计仍保持严重程度相关性(200词元/类时rho = -0.733),证明信号非词元数偏差所致。结果支持音位子空间分析作为稳健、无训练的病因感知构音障碍表征框架,具有跨语言轮廓稳定性与跨模型鲁棒性。

原文摘要 · Abstract (English)

We previously introduced a training-free method for dysarthria severity assessment based on d-prime separability of phonological feature subspaces in frozen self-supervised speech representations, validated on 890 speakers across 5 languages with HuBERT-base. Here, we scale the analysis to 3,374 speakers from 25 datasets spanning 12 languages and 5 aetiologies (Parkinson's disease, cerebral palsy, ALS, Down syndrome, and stroke), plus healthy controls, using 6 SSL backbones. We report three findings. First, aetiology-specific degradation profiles are distinguishable at the group level: 10 of 13 features yield large effect sizes (epsilon-squared > 0.14, Holm-corrected p < 0.001), with Parkinson's disease separable from the articulatory execution group at Cohen's d = 0.83; individual-level classification remains limited (22.6% macro F1). Second, profiles show cross-lingual profile-shape stability: cosine similarity of 5-dimensional consonant d-prime profiles exceeds 0.95 across the languages available for each aetiology. Absolute d-prime magnitudes are not cross-lingually calibrated, so the method supports language-independent phenotyping of degradation patterns but requires within-corpus calibration for absolute severity interpretation. Third, the method is architecture-independent: all 6 backbones produce monotonic severity gradients with inter-model agreement exceeding rho = 0.77. Fixed-token d-prime estimation preserves the severity correlation (rho = -0.733 at 200 tokens per class), confirming that the signal is not a token-count artefact. These results support phonological subspace analysis as a robust, training-free framework for aetiology-aware dysarthria characterisation, with evidence of cross-lingual profile-shape stability and cross-backbone robustness in the represented sample.

构音障碍自监督学习多病因分析跨语言研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。