用大模型生成符合指南的阿尔茨海默病诊断报告,支持多源数据、提升准确率与公平性。
AD-CARE: A Guideline-grounded, Modality-agnostic LLM Agent for Real-world Alzheimer's Disease Diagnosis with Multi-cohort Assessment, Fairness Analysis, and Reader Study
- 基于临床指南构建可适配多模态数据的智能诊断代理,无需补全缺失信息。
- 在10303例病例中达84.9%准确率,相对基线提升4.2%-13.7%,跨人群表现稳定。
- 显著降低不同种族和年龄群体的诊断偏差,帮助医生提速增效,适合临床部署。
阿尔茨海默病(AD)随人口老龄化日益成为全球健康挑战,及时准确的诊断对减轻个体与社会负担至关重要。然而,真实世界中的评估受限于多模态数据不完整、异质性强以及站点与人群差异。尽管大语言模型(LLMs)在生物医学领域展现潜力,但其在AD中的应用仍局限于回答特定问题,难以生成支持临床决策的综合性诊断报告。本文提出AD-CARE,一种模态无关的智能代理,能基于不完整、异质输入进行符合指南的诊断评估,无需补全缺失模态。通过动态调度专用诊断工具并嵌入临床指南到大模型推理中,生成透明、类报告式输出,契合真实临床流程。在涵盖10,303例病例的六个队列中,AD-CARE实现84.9%诊断准确率,相对基线提升4.2%-13.7%;各队列准确率保持稳健(80.4%-98.8%),持续优于所有基线。该框架有效降低种族与年龄亚组间的性能差异,四项指标的平均离散度分别下降21%-68%与28%-51%。在受控阅片研究中,显著提升神经科医生与放射科医生准确率(6%-11%),决策时间减少超一半。相比八种基础大模型,平均提升2.29%-10.66%且促进性能收敛。结果表明,AD-CARE是可扩展、可落地的多模态决策支持框架,适用于日常临床实践。
原文摘要 · Abstract (English)
Alzheimer's disease (AD) is a growing global health challenge as populations age, and timely, accurate diagnosis is essential to reduce individual and societal burden. However, real-world AD assessment is hampered by incomplete, heterogeneous multimodal data and variability across sites and patient demographics. Although large language models (LLMs) have shown promise in biomedicine, their use in AD has largely been confined to answering narrow, disease-specific questions rather than generating comprehensive diagnostic reports that support clinical decision-making. Here we expand LLM capabilities for clinical decision support by introducing AD-CARE, a modality-agnostic agent that performs guideline-grounded diagnostic assessment from incomplete, heterogeneous inputs without imputing missing modalities. By dynamically orchestrating specialized diagnostic tools and embedding clinical guidelines into LLM-driven reasoning, AD-CARE generates transparent, report-style outputs aligned with real-world clinical workflows. Across six cohorts comprising 10,303 cases, AD-CARE achieved 84.9% diagnostic accuracy, delivering 4.2%-13.7% relative improvements over baseline methods. Despite cohort-level differences, dataset-specific accuracies remain robust (80.4%-98.8%), and the agent consistently outperforms all baselines. AD-CARE reduced performance disparities across racial and age subgroups, decreasing the average dispersion of four metrics by 21%-68% and 28%-51%, respectively. In a controlled reader study, the agent improved neurologist and radiologist accuracy by 6%-11% and more than halved decision time. The framework yielded 2.29%-10.66% absolute gains over eight backbone LLMs and converges their performance. These results show that AD-CARE is a scalable, practically deployable framework that can be integrated into routine clinical workflows for multimodal decision support in AD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。