arXiv:2603.25196cs.CLcs.AI2026-03被引 3

首个系统评估大模型在多轮对话中识别和遵循临床指南能力的十年级基准。

A Decade-Scale Benchmark Evaluating LLMs' Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations

  • 构建跨24个专科的32,155条临床推荐数据集,生成多轮对话测试模型能力。
  • 大模型仅3.6%-29.7%能正确引用推荐来源,应用遵循率最高仅63.2%。
  • 首次揭示模型知其内容但难溯源与应用的鸿沟,适合医疗AI安全研究者参考。

临床实践指南(CPGs)对保障循证决策、改善患者结局至关重要。尽管大语言模型(LLMs)在医疗场景中日益应用,但其在对话中识别并遵循指南的能力仍不明确。为此,我们提出CPGBench,一个自动化框架,用于评估LLMs在多轮对话中检测和遵循临床指南的能力。我们从过去十年间9个国家/地区及2个国际组织收集了3,418份指南文档,涵盖24个专科,并从中提取32,155条临床推荐,包含发布机构、日期、国家、专科、推荐强度、证据等级等信息。每条推荐对应生成一段多轮对话,用于评估8个主流LLMs的检测与遵循能力。结果显示,71.1%-89.6%的推荐可被正确检测,但仅3.6%-29.7%的推荐标题能被正确引用,表明知晓内容与溯源之间存在差距。各模型遵循率在21.8%-63.2%之间,反映出知识掌握与实际应用之间的显著鸿沟。为验证自动分析的有效性,我们进一步开展涵盖56名不同专科医生的人工评估。据我们所知,CPGBench是首个系统揭示大模型在对话中未能检测或遵循哪些临床推荐的基准。鉴于每条推荐可能影响大量人群,且临床应用具有高度安全性,弥合这些差距对于大模型在真实临床实践中安全可靠部署至关重要。

原文摘要 · Abstract (English)

Clinical practice guidelines (CPGs) play a pivotal role in ensuring evidence-based decision-making and improving patient outcomes. While Large Language Models (LLMs) are increasingly deployed in healthcare scenarios, it is unclear to which extend LLMs could identify and adhere to CPGs during conversations. To address this gap, we introduce CPGBench, an automated framework benchmarking the clinical guideline detection and adherence capabilities of LLMs in multi-turn conversations. We collect 3,418 CPG documents from 9 countries/regions and 2 international organizations published in the last decade spanning across 24 specialties. From these documents, we extract 32,155 clinical recommendations with corresponding publication institute, date, country, specialty, recommendation strength, evidence level, etc. One multi-turn conversation is generated for each recommendation accordingly to evaluate the detection and adherence capabilities of 8 leading LLMs. We find that the 71.1%-89.6% recommendations can be correctly detected, while only 3.6%-29.7% corresponding titles can be correctly referenced, revealing the gap between knowing the guideline contents and where they come from. The adherence rates range from 21.8% to 63.2% in different models, indicating a large gap between knowing the guidelines and being able to apply them. To confirm the validity of our automatic analysis, we further conduct a comprehensive human evaluation involving 56 clinicians from different specialties. To our knowledge, CPGBench is the first benchmark systematically revealing which clinical recommendations LLMs fail to detect or adhere to during conversations. Given that each clinical recommendation may affect a large population and that clinical applications are inherently safety critical, addressing these gaps is crucial for the safe and responsible deployment of LLMs in real world clinical practice.

临床指南大模型评估医疗AI多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。