arXiv:2508.10777cs.AI2025-08

大模型有医学知识却不会推理,关键在于知识与推理分离

The Knowledge-Reasoning Dissociation: Fundamental Limitations of LLMs in Clinical Natural Language Inference

  • 用临床推理任务和知识验证探针分离知识获取与推理能力
  • 模型知识准确率达91.8%,但推理平均仅25%正确
  • 适合评估医疗等高风险领域AI的可靠性

大型语言模型常被认为通过数据和参数规模扩大可获得越来越结构化的通用内部表征。本文通过构建包含四种推理类型(因果归因、组合定位、认知验证、风险状态抽象)的临床试验自然语言推理基准,并结合针对性的底层知识与元级推理验证(GKMRV)探针,将事实获取失败与推理失败区分开来。在直接提示和思维链提示下评估六种主流LLM,模型在GKMRV任务上达到平均0.918的高准确率,但在主要推理任务上平均准确率仅为0.25。尽管推理错误率高,输出推理在样本间高度一致(平均0.87),表明存在系统性启发式策略。结果揭示当前大模型虽具备相关临床知识,但缺乏结构化、可组合的内部表示以可靠调用知识(如整合约束、权衡证据或模拟反事实)。通过GKMRV解耦知识与推理,使这种分离现象变得明确且可测量,为高风险领域大模型可靠性提供有效分析框架。

原文摘要 · Abstract (English)

Large language models are often assumed to acquire increasingly structured, generalizable internal representations simply by scaling data and parameters. We interrogate this assumption by introducing a Clinical Trial Natural Language Inference benchmark comprising four reasoning families, Causal Attribution, Compositional Grounding, Epistemic Verification, and Risk State Abstraction. Each item is paired with a targeted Ground Knowledge and Meta-Level Reasoning Verification (GKMRV) probe, allowing us to dissociate failures of factual access from failures of inference. We evaluate six contemporary LLMs under both direct and chain of thought prompting. Models achieve near-ceiling GKMRV accuracy (mean accuracy 0.918) yet perform poorly on the main reasoning tasks (mean accuracy 0.25). Despite low accuracy, output inferences are highly consistent across samples (mean 0.87), indicating a systematic application of underlying heuristics and shortcuts. These results reveal fundamental structural and representational limitations: current LLMs often possess the relevant clinical knowledge but lack the structured, composable internal representations needed to deploy it reliably (e.g., integrating constraints, weighing evidence, or simulating counterfactuals). Decoupling knowledge from reasoning with GKMRV makes this dissociation explicit and measurable, providing an effective framework for probing the reliability of LLMs in high-stakes domains.

大模型局限临床推理知识推理分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。