arXiv:2506.06603cs.LGcs.AI2025-06

用语音分析预测认知衰退,多模态方法更有效

CAtCh: Cognitive Assessment through Cookie Thief

  • 融合声学与情感特征的多模态方法提升预测效果
  • 声调与情绪相关声学特征显著优于语言模型特征
  • 适合临床早期筛查认知障碍的研究者参考

已有多种机器学习算法用于从自发言语中预测阿尔茨海默病及相关痴呆(ADRD)。然而,这些方法尚未应用于更广泛的认知障碍(CI)预测,而后者可能是ADRD的前兆和风险因素。本文评估了若干原本为ADRD预测设计的开源语音方法,以及来自多模态情感分析的方法,用于从患者音频中预测CI。结果表明,多模态方法在CI预测上优于单模态方法,声学方法优于语言学方法。具体而言,与基于BERT的语言特征和可解释语言特征相比,与情绪和语调相关的可解释声学特征表现显著更优。本研究所有代码均已公开于https://github.com/JTColonel/catch。

原文摘要 · Abstract (English)

Several machine learning algorithms have been developed for the prediction of Alzheimer's disease and related dementia (ADRD) from spontaneous speech. However, none of these algorithms have been translated for the prediction of broader cognitive impairment (CI), which in some cases is a precursor and risk factor of ADRD. In this paper, we evaluated several speech-based open-source methods originally proposed for the prediction of ADRD, as well as methods from multimodal sentiment analysis for the task of predicting CI from patient audio recordings. Results demonstrated that multimodal methods outperformed unimodal ones for CI prediction, and that acoustics-based approaches performed better than linguistics-based ones. Specifically, interpretable acoustic features relating to affect and prosody were found to significantly outperform BERT-based linguistic features and interpretable linguistic features, respectively. All the code developed for this study is available at https://github.com/JTColonel/catch.

认知评估语音分析多模态早期筛查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。