用语音的听觉和语言特征,自动识别阿尔茨海默病早期症状。
Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge
- 融合语音音频与语言特征,构建多模态分析模型。
- 语言特征组合在分类任务中达到最高F1值0.497。
- 语音嵌入模型对认知评分预测误差最小(RMSE=2.843)
阿尔茨海默病(AD)和轻度认知障碍(MCI)的早期检测对及时干预至关重要,但现有诊断方法仍资源密集且具侵入性。语音包含声学与语言双重维度,是认知衰退的潜在非侵入性生物标志物。本研究针对PROCESS挑战赛,提出一种机器学习框架,利用自发言语录音中的音频嵌入与语言特征。音频表示通过饼干盗窃任务的Whisper嵌入提取,语言特征则来自语义流畅性、音位流畅性及饼干盗窃图片描述的转录文本,涵盖代词使用、句法复杂度、填充词和从句结构等。分类模型用于区分健康对照组(HC)、MCI和AD人群,回归模型预测迷你精神状态检查(MMSE)得分。结果表明,基于语言特征拼接的投票集成模型在分类任务中表现最佳(F1 = 0.497),而基于Whisper嵌入的集成回归模型在预测中误差最低(RMSE = 2.843)。在PROCESS挑战赛中,我们的回归模型位列前列,分类模型居中,凸显语言与音频嵌入的互补优势。研究验证了多模态语音分析在可扩展、非侵入性认知评估中的潜力,并强调需结合任务特异性语言与声学标记以提升痴呆检测效果。
原文摘要 · Abstract (English)
Early detection of Alzheimer's Dementia (AD) and Mild Cognitive Impairment (MCI) is critical for timely intervention, yet current diagnostic approaches remain resource-intensive and invasive. Speech, encompassing both acoustic and linguistic dimensions, offers a promising non-invasive biomarker for cognitive decline. In this study, we present a machine learning framework for the PROCESS Challenge, leveraging both audio embeddings and linguistic features derived from spontaneous speech recordings. Audio representations were extracted using Whisper embeddings from the Cookie Theft description task, while linguistic features-spanning pronoun usage, syntactic complexity, filler words, and clause structure-were obtained from transcriptions across Semantic Fluency, Phonemic Fluency, and Cookie Theft picture description. Classification models aimed to distinguish between Healthy Controls (HC), MCI, and AD participants, while regression models predicted Mini-Mental State Examination (MMSE) scores. Results demonstrated that voted ensemble models trained on concatenated linguistic features achieved the best classification performance (F1 = 0.497), while Whisper embedding-based ensemble regressors yielded the lowest MMSE prediction error (RMSE = 2.843). Comparative evaluation within the PROCESS Challenge placed our models among the top submissions in regression task, and mid-range for classification, highlighting the complementary strengths of linguistic and audio embeddings. These findings reinforce the potential of multimodal speech-based approaches for scalable, non-invasive cognitive assessment and underline the importance of integrating task-specific linguistic and acoustic markers in dementia detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。