arXiv:2606.29876cs.CLcs.AI2026-06

用结构化图谱检测大模型诊断推理,发现准确但不一致。

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

论文配图:Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency
图 1 · 摘自论文原文
  • 构建临床推理图谱,用领域本体解析模型诊断过程
  • 50个病例中相似病例图谱相似度无显著差异
  • 适合评估医疗大模型推理质量,非仅看答案对错

现代大语言模型在复杂临床病例基准上达到60-70%的诊断准确率,但准确率无法区分稳定临床推理与模式匹配。本文提出临床推理图谱,基于领域本体(5种节点类型、7种边类型)从自由文本诊断轨迹中提取结构化图表示。针对5个模型在50例《新英格兰医学杂志》病案讨论案例及3种提示条件下生成的750条推理轨迹进行分析,检验诊断轨迹是否对临床相似病例表现出稳定的结构化推理模式(即诊断范式)。通过比较临床相似与不相似病例间的图谱相似度来量化该特征。在15组模型-提示组合中,组内与组间复合相似度几乎相等,且无一组通过多重检验校正;组件级分析显示残余内容信号远低于范式尺度。正确对与错误对的图谱相似度分别为0.488和0.484,表明图结构捕捉的是诊断准确率未反映的维度。结构化反思提示虽提升特征分析比例33%,但未提升跨案例一致性。结果表明存在诊断能力但缺乏范式级一致性,提示需以过程评估补充最终答案准确率。相关本体、提取管道、验证协议及提取的推理图谱与相似性数据已公开。

原文摘要 · Abstract (English)

Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable clinically-grounded reasoning from pattern matching. We introduce clinical reasoning graphs, structured graph representations extracted from free-text LLM diagnostic traces using a domain-grounded ontology with 5 node types and 7 edge types. We apply this pipeline to 750 traces from five LLMs across 50 New England Journal of Medicine Clinicopathological Conference cases and three prompt conditions, and test whether diagnostic traces show stable structured reasoning patterns, or diagnostic schemas, for clinically similar cases. We operationalize this as higher graph similarity among clinically similar cases than among clinically dissimilar ones. Across 15 model-condition comparisons, within-cluster and between-cluster composite similarity are nearly equal, and no comparison survives multiple-testing correction; a component-level analysis finds any residual content signal far below schema scale. Graph similarity is also nearly identical for pairs of models that are both correct (0.488) and both incorrect (0.484), suggesting that graph structure captures a dimension not reflected in diagnostic accuracy. Structured reflection prompting increases explicit discriminating-feature analysis within traces (+33%) but does not increase cross-case consistency. These results show diagnostic competence without schema-scale reasoning consistency, and indicate that final-answer accuracy should be complemented by process-level evaluation. We release the ontology, extraction pipeline, validation protocol, and the extracted reasoning graphs and similarity artifacts as resources for structured evaluation of LLM clinical reasoning.

医疗AI推理评估大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。