arXiv:2511.08132cs.AI2025-11被引 5

用多模态语音分析提前识别认知衰退,准确率超88%。

National Institute on Aging PREPARE Challenge: Early Detection of Cognitive Impairment Using Speech -- The SpeechCARE Solution

  • 融合声学、语言与人口统计信息的动态模型,提升检测精度。
  • 在阿尔茨海默病早期阶段识别中达AUC 0.90,F1值0.62。
  • 支持可解释性分析,适合医疗科研与临床辅助决策者使用。

阿尔茨海默病及相关痴呆症影响60岁以上成年人的五分之一,但超过半数的认知衰退患者未被诊断。语音评估有望实现早期发现,因发音运动规划缺陷会改变音高、语调等声学特征,而记忆和语言障碍则引发句法与语义错误。然而,传统基于手工特征或通用音频分类器的语音处理流程性能有限且泛化能力差。为此,我们提出SpeechCARE,一种基于预训练多语言声学与语言变换器模型的多模态语音处理管道,以捕捉与认知衰退相关的细微语音线索。受混合专家(MoE)范式启发,SpeechCARE采用动态融合架构,对声学、语言及人口统计输入进行加权整合,支持扩展其他模态(如社会因素、影像数据),并增强跨任务鲁棒性。其稳健预处理包括自动转录、大语言模型(LLM)异常检测与任务识别。通过SHAP可解释模块与LLM推理,突出各模态对决策的贡献。SpeechCARE在区分健康、轻度认知障碍(MCI)与阿尔茨海默病个体时达到AUC = 0.88、F1 = 0.72;在MCI检测中达AUC = 0.90、F1 = 0.62。偏差分析显示除80岁以上人群外,差异极小;缓解策略包括过采样与加权损失。未来工作将推进至真实护理场景(如VNS Health、哥伦比亚ADRC)及纽约市代表性不足群体的电子病历集成可解释系统。

原文摘要 · Abstract (English)

Alzheimer's disease and related dementias (ADRD) affect one in five adults over 60, yet more than half of individuals with cognitive decline remain undiagnosed. Speech-based assessments show promise for early detection, as phonetic motor planning deficits alter acoustic features (e.g., pitch, tone), while memory and language impairments lead to syntactic and semantic errors. However, conventional speech-processing pipelines with hand-crafted features or general-purpose audio classifiers often exhibit limited performance and generalizability. To address these limitations, we introduce SpeechCARE, a multimodal speech processing pipeline that leverages pretrained, multilingual acoustic and linguistic transformer models to capture subtle speech-related cues associated with cognitive impairment. Inspired by the Mixture of Experts (MoE) paradigm, SpeechCARE employs a dynamic fusion architecture that weights transformer-based acoustic, linguistic, and demographic inputs, allowing integration of additional modalities (e.g., social factors, imaging) and enhancing robustness across diverse tasks. Its robust preprocessing includes automatic transcription, large language model (LLM)-based anomaly detection, and task identification. A SHAP-based explainability module and LLM reasoning highlight each modality's contribution to decision-making. SpeechCARE achieved AUC = 0.88 and F1 = 0.72 for classifying cognitively healthy, MCI, and AD individuals, with AUC = 0.90 and F1 = 0.62 for MCI detection. Bias analysis showed minimal disparities, except for adults over 80. Mitigation techniques included oversampling and weighted loss. Future work includes deployment in real-world care settings (e.g., VNS Health, Columbia ADRC) and EHR-integrated explainability for underrepresented populations in New York City.

语音分析认知障碍多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。