用因果图评估医疗大模型决策,比传统方法更懂干预逻辑和风险。
Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot

- 构建可复现的因果知识图谱,让每个医学断言都有可追踪身份
- 四种不同接地条件测试下,集成式框架在因果推理上得分最高(F1=0.838)
- 适合研究医疗AI推理机制或想提升模型可解释性的团队
大型语言模型在医疗决策支持中日益受到关注,但现有评估仍偏向单一答案准确率,忽视对干预措施、作用机制、潜在危害、证据依据及不确定性等关键维度的推理能力。本文提出一种以图为中心的可复现评估框架,聚焦干预导向的医疗大模型行为,并在心血管疾病领域开展试点验证。框架包含四个部分:(i) 域内因果知识图谱,其中断言作为首等节点,具有溯源性与稳定标识符;(ii) 场景条件化的子图提取步骤,根据临床情景获取相关断言子图;(iii) 四种受控接地条件(未接地C1、知识图谱C2、因果图C3、集成式C4),变化上下文构成方式;(iv) 基于断言标识符的自动化评分流水线,单次运行即可计算干预准确率及其他指标。通过构建涵盖八类推理失败模式的平衡场景生成器并应用于心血管图谱,结果显示:在可解释且非冗余的维度上,指标能有效区分各条件;其中C4在因果边F1(0.838)、不良反应F1(0.833)、证据准确率(0.738)及无依据主张率(0.114)上表现最优,而C1虽干预准确率最高(0.948),但缺乏因果与证据根基。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。