arXiv:2602.13979cs.CL2026-02被引 1

用大模型推理链提升阿尔茨海默病诊断准确率与可解释性

Chain-of-Thought Reasoning with Large Language Models for Clinical Alzheimer's Disease Assessment and Diagnosis

  • 让大模型通过思维链逐步分析电子病历,生成诊断逻辑路径
  • 在多个临床痴呆评级任务中,F1分数比零样本基线最高提升15%
  • 适合需要可解释诊断的医疗AI研究者和临床决策支持系统开发者

阿尔茨海默病(AD)已成为全球高发的神经退行性疾病。传统诊断依赖医学影像和医生临床评估,耗时且对人力与医疗资源要求高。近年来,大语言模型(LLMs)开始应用于电子健康记录(EHR)的医疗场景,但在阿尔茨海默病评估中的应用仍有限,因其病因复杂、多因素交织,难以通过影像直接观察。本文提出利用LLM对患者临床EHR进行思维链(CoT)推理,而非直接微调模型进行分类。该方法生成显式的诊断推理路径,并基于结构化思维链输出预测结果。该流程不仅增强模型对复杂病理因素的诊断能力,也提升了不同疾病阶段诊断过程的可解释性。实验表明,所提出的基于思维链的诊断框架在多个临床痴呆量表(CDR)分级任务中显著提升稳定性和性能,相较零样本基线方法,最高实现15%的F1分数提升。

原文摘要 · Abstract (English)

Alzheimer's disease (AD) has become a prevalent neurodegenerative disease worldwide. Traditional diagnosis still relies heavily on medical imaging and clinical assessment by physicians, which is often time-consuming and resource-intensive in terms of both human expertise and healthcare resources. In recent years, large language models (LLMs) have been increasingly applied to the medical field using electronic health records (EHRs), yet their application in Alzheimer's disease assessment remains limited, particularly given that AD involves complex multifactorial etiologies that are difficult to observe directly through imaging modalities. In this work, we propose leveraging LLMs to perform Chain-of-Thought (CoT) reasoning on patients' clinical EHRs. Unlike direct fine-tuning of LLMs on EHR data for AD classification, our approach utilizes LLM-generated CoT reasoning paths to provide the model with explicit diagnostic rationale for AD assessment, followed by structured CoT-based predictions. This pipeline not only enhances the model's ability to diagnose intrinsically complex factors but also improves the interpretability of the prediction process across different stages of AD progression. Experimental results demonstrate that the proposed CoT-based diagnostic framework significantly enhances stability and diagnostic performance across multiple CDR grading tasks, achieving up to a 15% improvement in F1 score compared to the zero-shot baseline method.

阿尔茨海默病大模型推理可解释性电子病历

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。