大语言模型在真实临床中自主分诊仍不安全,因缺乏关键信息搜集能力。
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

- 模型优化于文本续写而非安全决策,难以识别高危但罕见的诊断。
- 在病史不全时,模型会漏掉关键警示信号,误判风险等级。
- 适合关注AI医疗安全、临床决策系统设计的研究者阅读。
大型语言模型虽已通过医学执照考试,在特定病例中可媲美医生的诊断推理能力,正被广泛应用于症状评估、治疗建议、文书撰写及规则提醒增强。然而,本文重点关注最具影响力的场景:完全由患者自述、未分化病情下无医生参与的自主分诊。目前尚无证据证明其安全性。问题不在医学知识,而在临床评估的准确性——一个以最可能文本延续为目标的模型,并非以安全行动为优化目标;当‘必须排除’的致命诊断概率极低时,模型却无法主动识别。安全分诊本质是代价不对称下的序列决策:一次致命遗漏远胜过多次误报。而关键信号往往患者未主动提供,且模型未被训练去主动寻找。核心缺陷在于不确定情况下的信息获取能力。在病史不完整时,模型难以表现出安全分诊所需的特征:拓宽鉴别诊断、主动寻求缺失红标、降低升级阈值、延迟判断直至信息充分、对高危害疾病保持警惕并及时升级。这些失效模式难以通过现有评估发现,因为当前测试多基于完整、精心筛选、带置信度过滤的模拟数据。此外,模型的助手式行为与认知偏差(如轻信、顺从、过度自信)可能加剧风险,若缺乏临床分诊逻辑约束。
原文摘要 · Abstract (English)
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。