发现顶级大模型在医学任务中过早下结论,影响诊断安全。
Quantifying and Mitigating Premature Closure in Frontier LLMs

- 定义并量化大模型在不确定时强行给出答案的现象
- 在500道医考题中,错误回答率高达55%-81%
- 适合医疗AI安全评估与提示工程研究者阅读
提前闭合指在信息不足时过早做出判断,是导致诊断失误的重要因素,但在大语言模型(LLMs)中尚未得到充分研究。本文将LLM的提前闭合定义为在不确定性下的不当承诺:在应澄清、回避或拒绝时仍提供答案、建议或临床指导。我们在结构化和开放性医学任务中评估了五款前沿LLM。在移除正确选项的MedQA(n=500)和AfriMed-QA(n=490)测试中,模型仍以高比例选择答案,基线错误动作率分别为55%-81%和53%-82%。在开放性评估中,模型在861道HealthBench问题中平均30%情况下给出不当回答,在191道医生设计的对抗性问题中达78%。安全导向提示可减少提前闭合,但残余失败依然存在,凸显需评估医学LLM何时不应作答。
原文摘要 · Abstract (English)
Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remains underexamined in large language models (LLMs). We define LLM premature closure as inappropriate commitment under uncertainty: providing an answer, recommendation, or clinical guidance when the safer response would be clarification, abstention, escalation, or refusal. We evaluated five frontier LLMs across structured and open-ended medical tasks. In MedQA (n = 500) and AfriMed-QA (n = 490) questions where the correct choice had been removed, models still selected an answer at high rates, with baseline false-action rates of 55-81% and 53-82%, respectively. In open-ended evaluation, models gave inappropriate answers on an average of 30% of 861 HealthBench questions and 78% of 191 physician-authored adversarial queries. Safety-oriented prompting reduced premature closure across models, but residual failure persisted, highlighting the need to evaluate whether medical LLMs know when not to answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。