强推理模型在医学难题中表现接近人类,更少受惯性思维干扰。
Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving
- 用对抗性医学问答数据集测试模型推理灵活性
- 顶级模型在关键题型上正确率达55%至70%,接近人类水平
- 适合医疗AI辅助决策、临床研究与模型评估者参考
大型语言模型(LLMs)在医学问答基准上已达到高准确率,但其在临床推理中的灵活性仍存争议。本文评估了来自OpenAI、Grok、Gemini、Claude和DeepSeek系列的推理模型在医学抽象与推理语料库(mARC)上的表现,该基准利用埃因斯滕效应诱导模型对习得启发式模式产生僵化依赖,而这些模式在特定情境下已不适用。结果显示,强推理模型比弱模型更少陷入埃因斯滕陷阱,在mARC上达到人类水平表现。对于医生常错的题目,前五名模型以高置信度正确回答了55%至70%的问题,表明这些模型可能比人类更不易受埃因斯滕效应影响。结果表明,强推理模型在医学推理中展现出更强的灵活性,表现已达人类水平。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their cognitive flexibility in clinical reasoning. We assessed reasoning models from the OpenAI, Grok, Gemini, Claude, and DeepSeek families on the medicine abstraction and reasoning corpus (mARC), an adversarial medical QA benchmark which utilizes the Einstellung effect to induce inflexible overreliance on learned heuristic patterns in contexts where they become suboptimal. We found that strong reasoning models avoided Einstellung-based traps more often than weaker reasoning models, achieving human-level performance on mARC. On questions most commonly missed by physicians, the top 5 performing models answered 55% to 70% correctly with high confidence, indicating that these models may be less susceptible than humans to Einstellung effects. Our results indicate that strong reasoning models demonstrate improved flexibility in medical reasoning, achieving performance on par with humans on mARC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。