测试大模型在医学诊断中是否能准确判断自身信心,发现其部分具备元认知敏感性。
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

- 设计心理物理实验框架,评估模型对证据强度的自信反应
- 准确率达93.5%,信心与证据质量正相关,但中等矛盾病例易误判
- 适合关注医疗AI可信度与错误模式的研究者使用
大型语言模型(LLMs)在医学领域应用日益广泛,其临床价值取决于答案准确性以及信心是否与证据质量及不确定性相匹配。我们开发了一个受控的、类心理物理学的临床基准,用于测试医学LLM在诊断选择和信心行为上的表现。研究聚焦于阿尔茨海默型神经认知障碍(AT-NCD)与抑郁相关认知障碍(DRCI)的区分。生成了45个合成病历,涵盖证据强度、冲突信息和缺失信息的变化,每种病历以三种提示变体呈现,共135次试验。在gpt-4.1-nano的初步测试中,所有试验均产生有效结构化输出。在强制选择试验中,诊断准确率为93.5%,平均信心为78.4%,AUROC2为0.876。信心随证据远离诊断边界而上升,缺失信息时下降,且在正确判断上信心高于错误判断(经证据强度和提示格式调整后)。结果表明模型具有部分元认知敏感性,而非完全无意义的信心。然而,错误集中在中等矛盾的AT-NCD病例中,模型更倾向于判断为DRCI,且信心高于实际准确率所支持水平。模型对比显示,信心质量应直接测量,而非仅从基准准确率或模型能力推断。本研究建立了一个可复现的框架,用于评估医学LLM的证据敏感性、元认知敏感性及局部校准失败。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。