用大模型分析画图描述,自动评估痴呆认知语言能力。
AI-based Cognitive-linguistic Features for Dementia Assessment in Picture Description

- 用大模型解析饼干盗窃图描述,量化七类认知语言特征。
- 最佳模型准确率达85%,能有效区分健康人与痴呆患者。
- 生成带解释的评分,临床医生认可度高,适合筛查工具开发。
画图描述能反映多种与认知语言能力相关的临床特征,但将其转化为可量化的指标仍具挑战,限制了可解释性和临床应用。我们针对饼干盗窃图片描述任务设计了七个评估维度,并利用大语言模型(LLMs)进行评分,生成严重程度分值及基于实例的解释。在对比的多个LLM中,Claude 3.5 Sonnet表现最优,其评分能显著区分认知受损者与健康对照组,在ADReSS数据集上达到85%的准确率。专家对Claude评分及解释的评价平均一致率达3.99/5。结果表明,大模型具备将临床构念操作化并生成可解释评估的能力,为开发可及的认知筛查工具提供了新路径。
原文摘要 · Abstract (English)
Picture descriptions provide valuable insights into several clinical constructs related to cognitive-linguistic abilities. However, operationalizing these constructs into quantitative measures remains challenging, limiting interpretability and clinical utility. We introduced seven constructs tailored to the Cookie Theft picture description task and prompted large language models (LLMs) to evaluate them, generating severity scores and example-based explanations. Among the examined LLMs, Claude 3.5 Sonnet performed the best, producing severity scores that significantly distinguish cognitively impaired individuals from healthy controls. The model achieves a high accuracy of 85% on the ADReSS dataset. Expert evaluation of Claude's scores and explanations yields a 3.99/5 average agreement. The findings demonstrate the potential of LLMs to operationalize clinical constructs and generate interpretable evaluations, offering a promising approach for accessible cognitive screening tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。