arXiv:2507.16836eess.AScs.LG2025-07被引 5

用稀疏自编码器解析帕金森语音模型,找到可解释的生物标志物。

From Black Box to Biomarker: Sparse Autoencoders for Interpreting Speech Models of Parkinson's Disease

  • 用基于掩码的稀疏自编码器提取语音模型中的可解释表征
  • 发现模型关注的低能量区谱流减少与帕金森病特征相关
  • 谱流与脑部苍白球体积关联,适合临床研究与疾病监测

语音有望成为帕金森病(PD)等神经疾病低成本、无创的生物标志物。尽管基于原始音频的深度学习系统能捕捉人工特征无法识别的细微信号,但其黑箱特性阻碍了临床应用。为此,本文采用稀疏自编码器(SAEs)从语音检测模型中提取可解释的内部表征。我们提出一种新型基于掩码的激活机制,使SAEs适用于小规模生物医学数据集,生成稀疏解耦的字典表征。这些字典条目与帕金森病语音的典型发音缺陷强相关,如模型注意力突出的低能量区域出现谱流降低和频谱平坦度升高。进一步发现,谱流与磁共振成像(MRI)测量的苍白球体积相关,表明SAEs有潜力揭示可用于疾病监测与诊断的临床生物标志物。

原文摘要 · Abstract (English)

Speech holds promise as a cost-effective and non-invasive biomarker for neurological conditions such as Parkinson's disease (PD). While deep learning systems trained on raw audio can find subtle signals not available from hand-crafted features, their black-box nature hinders clinical adoption. To address this, we apply sparse autoencoders (SAEs) to uncover interpretable internal representations from a speech-based PD detection system. We introduce a novel mask-based activation for adapting SAEs to small biomedical datasets, creating sparse disentangled dictionary representations. These dictionary entries are found to have strong associations with characteristic articulatory deficits in PD speech, such as reduced spectral flux and increased spectral flatness in the low-energy regions highlighted by the model attention. We further show that the spectral flux is related to volumetric measurements of the putamen from MRI scans, demonstrating the potential of SAEs to reveal clinically relevant biomarkers for disease monitoring and diagnosis.

语音分析帕金森病可解释性自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。