用临床指南约束大模型,提升癌症病历长期决策准确性。
CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records
- 将病历转为时间知识图谱,对齐指南路径做决策支持
- 在中英文数据集上显著优于基线模型,临床评估相关性高
- 适合医疗AI研发者、临床决策系统开发者使用
大语言模型在整合复杂、纵向的癌症电子病历方面潜力巨大,但面临三大挑战:难以处理长序列和碎片化病历进行准确时序分析;传统检索增强生成方法无法有效融入过程导向的临床指南,易产生临床幻觉;评估指标不可靠,阻碍系统验证。为此,我们提出CliCARE框架,通过将非结构化纵向EHR转化为患者特异的时间知识图谱(TKG),并将其与规范性指南知识图谱对齐,实现决策支持的可解释性与真实性。该方法生成高保真临床摘要与可操作建议。我们在大规模私有中文癌症数据集和公开英文MIMIC-IV数据集上验证,结果表明CliCARE显著优于主流长上下文LLM及知识图谱增强RAG方法。临床有效性通过严格评估协议验证,与肿瘤科医生评估高度相关。
原文摘要 · Abstract (English)
Large Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inability to effectively process the extensive length and fragmented nature of patient records for accurate temporal analysis; a heightened risk of clinical hallucination, as conventional grounding techniques such as Retrieval-Augmented Generation (RAG) do not adequately incorporate process-oriented clinical guidelines; and unreliable evaluation metrics that hinder the validation of AI systems in oncology. To address these issues, we propose CliCARE, a framework for Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records. The framework operates by transforming unstructured, longitudinal EHRs into patient-specific Temporal Knowledge Graphs (TKGs) to capture long-range dependencies, and then grounding the decision support process by aligning these real-world patient trajectories with a normative guideline knowledge graph. This approach provides oncologists with evidence-grounded decision support by generating a high-fidelity clinical summary and an actionable recommendation. We validated our framework using large-scale, longitudinal data from a private Chinese cancer dataset and the public English MIMIC-IV dataset. In these settings, CliCARE significantly outperforms baselines, including leading long-context LLMs and Knowledge Graph-enhanced RAG methods. The clinical validity of our results is supported by a robust evaluation protocol, which demonstrates a high correlation with assessments made by oncologists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。