用自监督学习提升帕金森语音诊断的跨语言泛化能力
Towards a Generalizable Speech Marker for Parkinson's Disease Diagnosis
- 基于HuBERT模型,通过自监督预训练提升语音特征提取能力
- 在多语言数据集上实现91.2%敏感度与92.1%特异度
- 适合需要非侵入性、低成本筛查的临床与研究场景
帕金森病(PD)是一种神经退行性疾病,早期即表现为语音异常。早期诊断对改善患者生活质量及提升疾病修饰疗法效果至关重要,但现有工具常错过这一关键窗口期。本文提出一种更具泛化能力的PD语音识别方法,结合领域自适应与自监督学习。该方法利用原始为语音识别训练的HuBERT模型,在老年群体的未标注语音数据上进行自监督微调,再在多语言数据集(英语、意大利语、西班牙语)中进行跨域微调。在四个公开的PD数据集上的评估表明,该方法平均敏感度达91.2%,平均特异度达92.1%。该方案可实现大人群中的客观一致评估,克服人工评估的变异性,提供一种无创、低成本且易获取的诊断选择。
原文摘要 · Abstract (English)
Parkinson's Disease (PD) is a neurodegenerative disorder characterized by motor symptoms, including altered voice production in the early stages. Early diagnosis is crucial not only to improve PD patients' quality of life but also to enhance the efficacy of potential disease-modifying therapies during early neurodegeneration, a window often missed by current diagnostic tools. In this paper, we propose a more generalizable approach to PD recognition through domain adaptation and self-supervised learning. We demonstrate the generalization capabilities of the proposed approach across diverse datasets in different languages. Our approach leverages HuBERT, a large deep neural network originally trained for speech recognition and further trains it on unlabeled speech data from a population that is similar to the target group, i.e., the elderly, in a self-supervised manner. The model is then fine-tuned and adapted for use across different datasets in multiple languages, including English, Italian, and Spanish. Evaluations on four publicly available PD datasets demonstrate the model's efficacy, achieving an average specificity of 92.1% and an average sensitivity of 91.2%. This method offers objective and consistent evaluations across large populations, addressing the variability inherent in human assessments and providing a non-invasive, cost-effective and accessible diagnostic option.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。