用知识图谱增强大模型推理,生成更准确的出院记录
Leaps Beyond the Seen: Reinforced Reasoning Augmented Generation for Clinical Notes
- 从医学知识图谱中检索推理路径,引导大模型生成
- 在输入信息少时仍能填补语义空白,生成更连贯内容
- 适合临床文本生成研究者和医疗AI开发者
临床笔记生成旨在基于患者病情和诊断过程生成自由文本摘要,出院指导是典型的长篇示例。尽管近期基于大语言模型(LLM)的方法在通用临床语料上预训练后展现出潜力,但在仅提供有限患者信息时难以生成长篇笔记。本文提出ReinRAG,一种基于预入院信息生成长篇出院指导的强化推理增强生成(RAG)方法。ReinRAG从医学知识图谱中检索推理路径,为大模型提供明确的语义引导。为弥补信息差距,我们提出分组式检索优化(GRO),通过分组归一化奖励提升检索质量,鼓励大模型进行更深层次的推理。在真实世界数据集上的全面实验表明,ReinRAG在临床有效性和自然语言生成指标上均优于基线。进一步分析显示,ReinRAG在输入稀疏场景下能有效填补语义空白,且检索到的推理路径帮助大模型聚焦关键证据、避免临床误判并保持推理连贯性。
原文摘要 · Abstract (English)
Clinical note generation aims to produce free-text summaries of a patient's condition and diagnostic process, with discharge instructions being a representative long-form example. While recent LLM-based methods pre-trained on general clinical corpora show promise in clinical text generation, they fall short in producing long-form notes from limited patient information. In this paper, we propose ReinRAG, a reinforced reasoning augmented generation (RAG) for long-form discharge instructions based on pre-admission information. ReinRAG retrieves reasoning paths from a medical knowledge graph to provide explicit semantic guidance to the LLM. To bridge the information gap, we propose group-based retriever optimization (GRO) which improves retrieval quality with group-normalized rewards, encouraging reasoning leaps for deeper inference by the LLM. Comprehensive experiments on the real-world dataset show that ReinRAG outperforms baselines in both clinical efficacy and natural language generation metrics. Further analysis reveals that ReinRAG fills semantic gaps in sparse input scenarios, and retrieved reasoning paths help LLMs avoid clinical misinterpretation by focusing on key evidence and following coherent reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。