arXiv:2605.26567cs.AI2026-05被引 2

将临床指南转化为可执行决策逻辑,提升医学大模型推理能力

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

论文配图:MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
图 1 · 摘自论文原文
  • 将指南转化为可执行决策逻辑,生成真实与反事实问答数据
  • 在4个基准上平均准确率提升10.28%
  • 更适合临床医生使用,推理过程更符合专业习惯

临床实践指南(CPGs)包含基于证据的决策逻辑,医生通过评估患者变量、条件标准和推荐规则来应用。现有方法多将指南作为自由文本训练数据或检索源,未能充分利用其程序化决策结构。为此,我们提出一种指南衍生的训练流程,将指南推荐转换为可执行的临床决策逻辑,并据此生成真实与反事实的问答数据。这些数据使模型既学习指南支持的决策,也理解不同患者条件下决策的变化。在生成数据上微调医学大模型后得到MedGuideX。在四个临床推理基准上,MedGuideX平均准确率相对提升10.28%。医生评估显示,MedGuideX更准确还原医生撰写的推理步骤,且在忠实性、有效性、完整性与清晰度方面更受医生青睐。结果表明,可执行的指南决策逻辑可转化为构建可靠医学大模型的可扩展监督信号。

原文摘要 · Abstract (English)

Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, and recommendation rules. However, existing methods often use CPGs as free-text training data or retrieval sources, underutilizing their procedural decision structure. To better exploit this structure, we introduce a guideline-derived training pipeline that transforms CPG recommendations into executable clinical decision logic and uses it to generate factual and counterfactual question-answering data. Theses data teach models both guideline-supported decisions and how decisions change under different patient conditions. Post-training a medical LLM on the generated data yields MedGuideX. Across four clinical reasoning benchmarks, MedGuideX achieves a 10.28% relative improvement in average accuracy. Physician evaluation further shows that MedGuideX better recovers clinician authored reasoning steps and produces physician-preferred rationales in faithfulness, validity, completeness, and clarity. Overall, our results show that executable decision logic from CPGs can be transformed into scalable supervision for building reliable medical LLMs.

医学大模型临床推理决策逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。