arXiv:2511.20680cs.CLcs.AI2025-11被引 1

大模型推理有偏易出错,临床决策需警惕虚假可信

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

  • 构建三级错误分类体系,关联计算失误与认知偏差
  • 23%的临床解读存在推理错误,高级别疾病中更易导致危险建议
  • 现有自动评估无法区分错误类型,需人工校验

尽管大型语言模型在临床基准测试中表现优异,但可能通过错误推理得出正确结论,这种安全风险在以准确率为核心的评估中无法被发现。本研究基于CORAL数据集中的乳腺癌和胰腺癌病历,对GPT-4的链式思维输出进行600条标注,构建了映射至认知偏差框架的三级推理错误分类体系。该体系在822份前列腺癌会诊记录(涵盖局部至转移性病变)中验证,模拟了信息提取、分析与临床建议任务。结果显示,23%的解释存在推理错误,确认偏差和锚定偏差最为常见;推理失败与指南不符及潜在有害建议显著相关,尤其在晚期疾病管理中。使用先进语言模型的自动化评估器虽能检测错误存在,却无法可靠识别子类型。研究揭示:大模型可能提供流畅但临床不安全的建议。该分类体系为临床部署前评估和提升推理可靠性提供了可推广的框架。

原文摘要 · Abstract (English)

Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.

大模型推理临床决策认知偏差医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。