测试大模型在临床推理中的僵化思维缺陷,发现其易出错且盲目自信。
Limitations of Large Language Models in Clinical Problem-Solving Arising from Inflexible Reasoning
- 设计新数据集M-ARC,用心理定势测试模型是否死记硬背而非灵活思考。
- 顶尖模型在复杂临床场景中准确率远低于医生,错误率达40%以上。
- 适合关注AI医疗安全与可解释性的研究人员和临床决策支持开发者。
大型语言模型(LLMs)在医学问答基准上已达到人类水平的准确率。然而,它们在处理开放式临床情境时的局限性日益显现,引发对其推理鲁棒性和泛化能力的担忧。为探究大模型在临床问题解决中的潜在失败模式,我们提出了医学抽象与推理语料库(M-ARC)。M-ARC通过设计利用“心理定势”效应的临床场景,测试模型是否因训练数据中的归纳偏差而陷入僵化模式匹配,而非进行灵活推理。实验发现,包括当前最先进的o1和Gemini模型在内的多种大模型,在M-ARC上的表现显著低于医生,常出现常识性医学推理缺失及幻觉现象。此外,不确定性分析表明,尽管准确率有限,这些模型仍表现出过度自信。M-ARC揭示的失败模式凸显了在临床环境中部署此类模型时需保持审慎。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have attained human-level accuracy on medical question-answer (QA) benchmarks. However, their limitations in navigating open-ended clinical scenarios have recently been shown, raising concerns about the robustness and generalizability of LLM reasoning across diverse, real-world medical tasks. To probe potential LLM failure modes in clinical problem-solving, we present the medical abstraction and reasoning corpus (M-ARC). M-ARC assesses clinical reasoning through scenarios designed to exploit the Einstellung effect -- the fixation of thought arising from prior experience, targeting LLM inductive biases toward inflexible pattern matching from their training data rather than engaging in flexible reasoning. We find that LLMs, including current state-of-the-art o1 and Gemini models, perform poorly compared to physicians on M-ARC, often demonstrating lack of commonsense medical reasoning and a propensity to hallucinate. In addition, uncertainty estimation analyses indicate that LLMs exhibit overconfidence in their answers, despite their limited accuracy. The failure modes revealed by M-ARC in LLM medical reasoning underscore the need to exercise caution when deploying these models in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。