arXiv:2506.21615cs.CLcs.AI2025-06被引 2

用临床指南增强诊断模型,避免幻觉且更符合医生决策逻辑。

Refine Medical Diagnosis Using Generation Augmented Retrieval and Clinical Practice Guidelines

  • 用病历和大模型预测生成查询,精准检索临床指南片段。
  • 在高血压诊断上,检索精度和指南符合度均优于传统方法。
  • 适合希望提升AI诊断可信度的医疗机构使用。

当前医学语言模型多基于ICD编码进行诊断预测,但此类标签无法体现医生决策中的上下文与推理过程。临床决策依赖患者多源数据及权威临床实践指南(CPGs)。本文提出GARMLE-G框架,通过直接检索权威指南内容,实现无幻觉输出。该框架将大模型预测与电子健康记录(EHR)融合生成语义丰富查询,利用嵌入相似性检索相关指南片段,并将指南内容与模型输出融合生成临床一致推荐。针对高血压诊断的原型系统在多个指标上优于基于RAG的基线,具备高检索精度、强语义相关性与良好指南遵循度,同时保持轻量化结构,适合本地化医疗部署。本工作为医学大模型提供了一种可扩展、低成本、无幻觉的证据驱动方法,具备广泛临床应用潜力。

原文摘要 · Abstract (English)

Current medical language models, adapted from large language models (LLMs), typically predict ICD code-based diagnosis from electronic health records (EHRs) because these labels are readily available. However, ICD codes do not capture the nuanced, context-rich reasoning clinicians use for diagnosis. Clinicians synthesize diverse patient data and reference clinical practice guidelines (CPGs) to make evidence-based decisions. This misalignment limits the clinical utility of existing models. We introduce GARMLE-G, a Generation-Augmented Retrieval framework that grounds medical language model outputs in authoritative CPGs. Unlike conventional Retrieval-Augmented Generation based approaches, GARMLE-G enables hallucination-free outputs by directly retrieving authoritative guideline content without relying on model-generated text. It (1) integrates LLM predictions with EHR data to create semantically rich queries, (2) retrieves relevant CPG knowledge snippets via embedding similarity, and (3) fuses guideline content with model output to generate clinically aligned recommendations. A prototype system for hypertension diagnosis was developed and evaluated on multiple metrics, demonstrating superior retrieval precision, semantic relevance, and clinical guideline adherence compared to RAG-based baselines, while maintaining a lightweight architecture suitable for localized healthcare deployment. This work provides a scalable, low-cost, and hallucination-free method for grounding medical language models in evidence-based clinical practice, with strong potential for broader clinical deployment.

医学AI临床指南检索增强大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。