arXiv:2607.10275cs.AIcs.CL2026-07被引 1

大模型在临床推理中常因信息搜寻不足而误诊,而非知识不够。

Information-seeking failures of large language models in agentic clinical reasoning

论文配图:Information-seeking failures of large language models in agentic clinical reasoning
图 1 · 摘自论文原文
  • 让模型主动分三轮申请临床数据,模拟真实诊断流程。
  • 最佳模型准确率仅68%,最后一轮数据请求率暴跌至26%。
  • 模型看似逻辑清晰,实则受认知偏差影响,适合临床研究者参考。

大型语言模型在医学知识测试中表现优异,但临床推理需要在不确定性下主动决策该调查什么。我们构建了一个血液肿瘤领域的代理式评估框架,要求模型在做出诊断和治疗方案前,分三个阶段主动请求临床数据。32个前沿模型中,表现最好的准确率仅为68%。信息利用效率(即实际请求的数据占比)是诊断准确性的最强预测因子(R = 0.69, P < 0.001),但在最后一轮,该比例从57%骤降至26%,导致对治疗选择至关重要的分子与细胞遗传学数据未被考察。尽管推理过程在临床推理评分中达到91%高于阈值,但其与准确性脱钩,揭示了局部逻辑连贯与全局正确结论之间的差距。错误分析表明,搜索满足、锚定效应和过早闭合是主要失败模式,这些正是双过程模型下新手医生的认知偏差。结果表明,当前模型在临床肿瘤学中的主要瓶颈并非医学知识不足,而是面对不确定性时系统性地缺乏信息搜寻能力。

原文摘要 · Abstract (English)

Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematologic oncology in which models must proactively request clinical data across three sequential rounds before committing to a diagnosis and treatment plan. Across 32 frontier models, the best achieved only 68% overall accuracy. Information utilization, the fraction of available data actually requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P < 0.001), yet utilization collapsed from 57% to 26% in the final round, leaving molecular and cytogenetic data critical for treatment selection unexamined. Reasoning traces scored high on a clinical reasoning rubric (91% above threshold) but decorrelated from accuracy, revealing a gap between locally coherent rationales and globally correct conclusions. Error analysis identified search satisficing, anchoring and premature closure as the dominant failure modes, the same cognitive biases that characterize novice clinicians under dual-process models of diagnostic reasoning. These findings demonstrate that the primary limitation of current models in clinical oncology is not insufficient medical knowledge but a systematic failure of information-seeking under uncertainty.

临床推理大模型认知偏差信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。