首个针对阿尔茨海默病的LLM评测基准,覆盖临床与照护场景。
ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias
- 构建融合7个医学基准的统一问答集,评估临床知识掌握能力。
- 引入149道照护场景新题,填补现有评测中实际照护情境空白。
- 发现顶尖模型虽准确率超90%,但推理质量不稳定,需加强领域适配。
大型语言模型(LLMs)在医疗领域展现出巨大潜力,但现有评测基准对阿尔茨海默病及相关痴呆症(ADRD)的覆盖有限。为此,我们提出ADRD-Bench,一个初步的ADRD专用LLM评测基准。该基准包含两部分:1)ADRD Unified QA,整合自七个权威医学基准的1,438个问题,用于统一评估临床知识;2)ADRD Caregiving QA,基于国家级大型临床试验支持的脑健康项目生成的149个照护相关问题,弥补现有基准在实际照护情境上的缺失。我们在ADRD-Bench上评估了36个前沿LLM。结果显示,开源通用模型、开源医学模型及顶尖闭源通用模型的准确率分别为0.63–0.93(均值0.77,标准差0.09)、0.47–0.93(均值0.81,标准差0.14)、0.83–0.93(均值0.90,标准差0.03)。尽管顶级模型准确率超过0.9,案例分析显示其推理质量和稳定性不一致,凸显出亟需基于日常照护数据提升模型在该领域的知识与推理能力。完整数据集可访问 https://github.com/IIRL-ND/ADRD-Bench。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown great potential for healthcare applications. However, existing evaluation benchmarks provide minimal coverage of Alzheimer's Disease and Related Dementias (ADRD). To address this gap, we introduce ADRD-Bench, a preliminary ADRD-specific LLM benchmark. ADRD-Bench has two components: 1) ADRD Unified QA, a synthesis of 1,438 questions consolidated from seven established medical benchmarks, providing a unified assessment of clinical knowledge; and 2) ADRD Caregiving QA, a novel set of 149 questions derived from a nationally adopted, large clinical trials supported brain health management program, mitigating the lack of practical caregiving context in existing benchmarks. We evaluated 36 state-of-the-art LLMs on the proposed ADRD-Bench. Results showed that the accuracy of open-weight general models, open-weight medical models, and frontier closed-source general models ranged from 0.63 to 0.93 (mean: 0.77; std: 0.09), 0.47 to 0.93 (mean: 0.81; std: 0.14), and 0.83 to 0.93 (mean: 0.90; std: 0.03), respectively. While top-tier models achieved high accuracies (>0.9), case studies revealed inconsistent reasoning quality and stability, highlighting a critical need for domain-specific improvement to enhance LLMs' knowledge and reasoning grounded in daily caregiving data. The entire dataset is available at https://github.com/IIRL-ND/ADRD-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。