arXiv:2510.25628cs.CL2025-10被引 14

构建医疗记录推理模型,提升临床决策智能化水平

EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis

  • 基于思维图框架自动生成30万条医疗推理数据
  • 720亿参数模型在42项任务上超越GPT-4o超30分
  • 专为电子病历设计,适合临床研究与医疗AI开发

电子健康记录(EHR)包含丰富但复杂的信息,其自动化分析对临床决策至关重要。尽管大语言模型在临床流程中取得进展,但其在EHR分析中的能力仍受限于任务覆盖窄和缺乏医疗推理能力。本文提出EHR-Ins,一个大规模、全面的医疗推理指令数据集,包含30万条高质量推理案例和400万条非推理案例,涵盖42类EHR任务。其核心创新是思维图驱动的生成框架,可规模化生成高质量推理数据。基于此,我们开发了系列高达720亿参数的推理增强型大模型EHR-R1,通过多阶段训练(领域适配、推理增强、强化学习),系统性获取领域知识与多样推理能力,实现精准可靠的EHR分析。最后,我们构建EHR-Bench,从MIMIC-IV中精选42项任务,用于全面评估推理与预测能力。实验表明,EHR-R1持续优于现有商用及开源模型(包括DeepSeek-V3和GPT-4o),在MIMIC-Bench上超越GPT-4o超过30分,在EHRSHOT上零样本AUROC高出10%。综合来看,EHR-Ins、EHR-R1与EHR-Bench显著推动了更可靠、更具临床意义的EHR分析发展。

原文摘要 · Abstract (English)

Electronic Health Records (EHRs) contain rich yet complex information, and their automated analysis is critical for clinical decision-making. Despite recent advances of large language models (LLMs) in clinical workflows, their ability to analyze EHRs remains limited due to narrow task coverage and lack of EHR-oriented reasoning capabilities. This paper aims to bridge the gap, specifically, we present EHR-Ins, a large-scale, comprehensive EHR reasoning instruction dataset, comprising 300k high-quality reasoning cases and 4M non-reasoning cases across 42 distinct EHR tasks. Its core innovation is a thinking-graph-driven framework that enables to generate high-quality reasoning data at scale. Based on it, we develop EHR-R1, a series of reasoning-enhanced LLMs with up to 72B parameters tailored for EHR analysis. Through a multi-stage training paradigm, including domain adaptation, reasoning enhancement, and reinforcement learning, EHR-R1 systematically acquires domain knowledge and diverse reasoning capabilities, enabling accurate and robust EHR analysis. Lastly, we introduce EHR-Bench, a new benchmark curated from MIMIC-IV, spanning 42 tasks, to comprehensively assess reasoning and prediction across EHR scenarios. In experiments, we show that the resulting EHR-R1 consistently outperforms state-of-the-art commercial and open-source LLMs (including DeepSeek-V3 and GPT-4o), surpassing GPT-4o by over 30 points on MIMIC-Bench and achieving a 10\% higher zero-shot AUROC on EHRSHOT. Collectively, EHR-Ins, EHR-R1, and EHR-Bench have significantly advanced the development for more reliable and clinically relevant EHR analysis.

医疗AI大模型推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。