arXiv:2512.21127cs.AI2025-12被引 6

首次在真实医疗数据中评估大模型药物安全审查,发现其易因理解偏差出错。

A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care

  • 用真实英国基层医疗数据测试大模型药物安全检测能力
  • 识别问题敏感度100%,但正确处理全案仅46.9%
  • 主要失败源于对临床情境理解不足,非知识缺失

大型语言模型(LLMs)在医学基准测试中常达到或超过临床医生水平,但极少在真实临床数据上评估,也未深入分析关键失败行为。我们首次在英国国家医疗服务体系(NHS)柴郡与默西赛德地区的真实基层医疗数据上,对基于大模型的药物安全审查系统进行评估,涵盖不同临床复杂度和用药风险的患者群体。研究基于2,125,549名成年人的电子健康记录(EHR),经筛选后保留277例患者,由专家临床医生评审系统识别的问题及建议干预措施。主模型在识别问题存在性方面表现优异(敏感性100% [95% CI 98.2–100],特异性83.1% [95% CI 72.7–90.1]),但在完整准确识别所有问题与干预措施方面仅达46.9% [95% CI 41.1–52.8]。失败分析揭示,主要失效机制是上下文推理错误,而非缺乏药物知识,共识别出五类典型模式:对不确定性的过度自信、机械套用指南忽略个体情况、误解实际医疗运作方式、事实性错误、流程盲视。这些模式在不同复杂度和人口统计学分组中均持续存在,并跨越多个前沿模型与配置。研究提供45个详细案例,全面覆盖所有失败情形。结果表明,大模型临床应用前需解决此类缺陷,并呼吁开展更大规模前瞻性评估与更深层行为研究。

原文摘要 · Abstract (English)

Large language models (LLMs) often match or exceed clinician-level performance on medical benchmarks, yet very few are evaluated on real clinical data or examined beyond headline metrics. We present, to our knowledge, the first evaluation of an LLM-based medication safety review system on real NHS primary care data, with detailed characterisation of key failure behaviours across varying levels of clinical complexity. In a retrospective study using a population-scale EHR spanning 2,125,549 adults in NHS Cheshire and Merseyside, we strategically sampled patients to capture a broad range of clinical complexity and medication safety risk, yielding 277 patients after data-quality exclusions. An expert clinician reviewed these patients and graded system-identified issues and proposed interventions. Our primary LLM system showed strong performance in recognising when a clinical issue is present (sensitivity 100\% [95\% CI 98.2--100], specificity 83.1\% [95\% CI 72.7--90.1]), yet correctly identified all issues and interventions in only 46.9\% [95\% CI 41.1--52.8] of patients. Failure analysis reveals that, in this setting, the dominant failure mechanism is contextual reasoning rather than missing medication knowledge, with five primary patterns: overconfidence in uncertainty, applying standard guidelines without adjusting for patient context, misunderstanding how healthcare is delivered in practice, factual errors, and process blindness. These patterns persisted across patient complexity and demographic strata, and across a range of state-of-the-art models and configurations. We provide 45 detailed vignettes that comprehensively cover all identified failure cases. This work highlights shortcomings that must be addressed before LLM-based clinical AI can be safely deployed. It also begs larger-scale, prospective evaluations and deeper study of LLM behaviours in clinical contexts.

药物安全大模型临床评估真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。