arXiv:2604.10455cs.CL2026-04KDD

用深度模型引导大模型,提升医疗诊断预测准确性。

EviCare: Enhancing Diagnosis Prediction with Deep Model-Guided Evidence for In-Context Reasoning

  • 先用深度模型筛选候选诊断,再构建证据关系,生成引导提示。
  • 在真实数据集上平均提升20.65%精度与准确率,新病种预测提升30.97%。
  • 适合需要高可靠性、早期发现罕见病的临床辅助系统使用。

大语言模型在电子健康记录(EHR)诊断预测中取得进展,但现有方法易过度依赖历史诊断,忽视关键的新发疾病。为此,我们提出EviCare,一种将深度模型引导融入上下文推理的框架。该框架不直接以原始EHR输入提示大模型,而是通过(1)深度模型候选诊断筛选,(2)基于集合的证据优先级排序,(3)为新诊断构建关系性证据,生成自适应上下文提示,引导大模型进行精准可解释的推理。在两个真实世界EHR基准数据集(MIMIC-III和MIMIC-IV)上的实验表明,EviCare在精度与准确率指标上平均优于纯大模型与纯深度模型基线20.65%。在挑战性的新诊断预测任务中,性能平均提升达30.97%。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have enabled promising progress in diagnosis prediction from electronic health records (EHRs). However, existing LLM-based approaches tend to overfit to historically observed diagnoses, often overlooking novel yet clinically important conditions that are critical for early intervention. To address this, we propose EviCare, an in-context reasoning framework that integrates deep model guidance into LLM-based diagnosis prediction. Rather than prompting LLMs directly with raw EHR inputs, EviCare performs (1) deep model inference for candidate selection, (2) evidential prioritization for set-based EHRs, and (3) relational evidence construction for novel diagnosis prediction. These signals are then composed into an adaptive in-context prompt to guide LLM reasoning in an accurate and interpretable manner. Extensive experiments on two real-world EHR benchmarks (MIMIC-III and MIMIC-IV) demonstrate that EviCare achieves significant performance gains, which consistently outperforms both LLM-only and deep model-only baselines by an average of 20.65\% across precision and accuracy metrics. The improvements are particularly notable in challenging novel diagnosis prediction, yielding average improvements of 30.97\%.

医疗AI诊断预测大模型证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。