评测大模型遵循临床指南路径的能力,发现现有模型仍难可靠支持医疗决策。
Benchmarking Clinical Decision Pathway Adherence in Large Language Models

- 基于指南构建4.2万例临床案例,评估模型生成合规诊疗路径能力。
- 16个主流大模型在路径一致性上表现不佳,平均仅58%符合指南要求。
- 适合医疗AI研发者和临床决策系统开发者参考,推动指南遵循性评估。
遵循由临床实践指南定义的临床决策路径(CDPs)对于安全可靠的医疗决策至关重要。然而,现有的医学大语言模型(LLM)基准主要评估最终答案的准确性,对模型遵循指南的能力评估有限。为填补这一空白,我们提出MEGA-CDP,一个用于评估医学LLM能否基于给定指南生成符合指南的临床决策路径的基准。MEGA-CDP通过指南到病例的转化流程,从2,274份英文和中文临床实践指南中构建,生成了包含42,353个明确参考路径的临床案例。该基准支持单轮情景题和多轮交互两种设置,并引入面向路径的一致性评估框架。在16个代表性医学大模型上的实验表明,当前模型实现可靠临床决策支持依然困难,凸显了面向路径评估的必要性,也验证了MEGA-CDP在提升医学大模型指南遵循性方面的价值。
原文摘要 · Abstract (English)
Following clinical decision pathways (CDPs) defined by clinical practice guidelines is essential for safe and reliable medical decision-making. However, existing medical large language model (LLM) benchmarks mainly evaluate final-answer accuracy, providing limited evaluation of models' ability to adhere to guidelines. To address this gap, we introduce MEGA-CDP, a benchmark for evaluating whether medical LLMs can generate guideline-adherent CDPs using provided guidelines as references. MEGA-CDP is constructed from 2,274 English and Chinese clinical practice guidelines through a guideline-to-case pipeline, yielding 42,353 clinical cases with explicit reference CDPs. It supports both single-turn vignette and multi-turn interactive settings, and introduces a CDP-oriented evaluation framework for measuring pathway consistency. Experiments on 16 representative LLMs show that reliable clinical decision support remains challenging for current models, demonstrating the need for CDP-oriented evaluation and the value of MEGA-CDP for advancing guideline adherence in medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。