arXiv:2601.03475cs.AI2026-01被引 2

将临床指南转化为大模型可执行的决策支持系统,提升医疗AI的准确性与可解释性。

CPGPrompt: Translating Clinical Guidelines into LLM-Executable Decision Support

  • 将文本式临床指南转为结构化决策树,由大模型动态推理患者诊疗路径。
  • 在专科转诊任务中表现优异(F1: 0.85–1.00,召回率1.00),但多路径分类受指南结构影响。
  • 适用于需精准遵循指南的医疗AI场景,尤其适合实验室指标明确的疾病领域。

临床实践指南(CPGs)为患者治疗提供循证建议,但将其融入人工智能仍具挑战。以往基于规则的方法存在可解释性差、遵循不一致和领域局限等问题。为此,我们提出CPGPrompt,一种自动提示系统,将叙述性临床指南转化为大语言模型(LLMs)可执行的形式。该框架将指南转换为结构化决策树,并利用大模型动态导航以评估患者病例。我们在头痛、腰痛和前列腺癌三个领域生成合成病例,分为四类测试不同决策场景。系统在二分类专科转诊任务中表现稳定(F1: 0.85–1.00),召回率高达1.00 ± 0.00;而在多类别路径分类任务中表现下降,领域差异明显:头痛(F1: 0.47)、腰痛(F1: 0.72)、前列腺癌(F1: 0.77)。性能差异源于各指南结构特征:头痛指南对否定句处理困难,腰痛指南需时间推理,而前列腺癌因有可量化的实验室检查,决策更可靠。

原文摘要 · Abstract (English)

Clinical practice guidelines (CPGs) provide evidence-based recommendations for patient care; however, integrating them into Artificial Intelligence (AI) remains challenging. Previous approaches, such as rule-based systems, face significant limitations, including poor interpretability, inconsistent adherence to guidelines, and narrow domain applicability. To address this, we develop and validate CPGPrompt, an auto-prompting system that converts narrative clinical guidelines into large language models (LLMs). Our framework translates CPGs into structured decision trees and utilizes an LLM to dynamically navigate them for patient case evaluation. Synthetic vignettes were generated across three domains (headache, lower back pain, and prostate cancer) and distributed into four categories to test different decision scenarios. System performance was assessed on both binary specialty-referral decisions and fine-grained pathway-classification tasks. The binary specialty referral classification achieved consistently strong performance across all domains (F1: 0.85-1.00), with high recall (1.00 $\pm$ 0.00). In contrast, multi-class pathway assignment showed reduced performance, with domain-specific variations: headache (F1: 0.47), lower back pain (F1: 0.72), and prostate cancer (F1: 0.77). Domain-specific performance differences reflected the structure of each guideline. The headache guideline highlighted challenges with negation handling. The lower back pain guideline required temporal reasoning. In contrast, prostate cancer pathways benefited from quantifiable laboratory tests, resulting in more reliable decision-making.

医疗AI决策支持大模型应用指南转化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。