arXiv:2502.08862eess.AS2025-02被引 9

用语音多模态分析预测认知衰退,准确区分轻度痴呆与阿尔茨海默病。

Predicting Cognitive Decline: A Multimodal AI Approach to Dementia Screening from Speech

  • 融合声学特征与Whisper/RoBERTa语音文本嵌入,提升识别能力。
  • 回归任务RMSE达2.77,分类任务宏F1达0.58,性能优于基线。
  • 提出两阶段分类框架,更好区分轻度认知障碍与痴呆患者。

近年来,仅通过患者语音录音检测早期痴呆取得进展。本研究参与PROCESS挑战赛,利用临床访谈音频预测受试者为健康对照、轻度认知障碍(MCI)或痴呆,并回归其迷你精神状态检查(MMSE)得分。方法结合声学特征(eGeMAPS与韵律)及Whisper和RoBERTa模型的嵌入表示,在回归任务中实现RMSE 2.7666,分类任务宏F1得分为0.5774。此外,设计新颖的两阶段分类架构以更好区分MCI与痴呆。该方法在测试集上排名第七(回归)和第十一(分类),共37支队伍,优于基线表现。

原文摘要 · Abstract (English)

Recent progress has been made in detecting early stage dementia entirely through recordings of patient speech. Multimodal speech analysis methods were applied to the PROCESS challenge, which requires participants to use audio recordings of clinical interviews to predict patients as healthy control, mild cognitive impairment (MCI), or dementia and regress the patient's Mini-Mental State Exam (MMSE) scores. The approach implemented in this work combines acoustic features (eGeMAPS and Prosody) with embeddings from Whisper and RoBERTa models, achieving competitive results in both regression (RMSE: 2.7666) and classification (Macro-F1 score: 0.5774) tasks. Additionally, a novel two-tiered classification setup is utilized to better differentiate between MCI and dementia. Our approach achieved strong results on the test set, ranking seventh on regression and eleventh on classification out of thirty-seven teams, exceeding the baseline results.

认知衰退语音分析多模态AI痴呆筛查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。