用病历文本学习重症治疗的奖励函数,提升治疗效果评估
Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment

- 从出院小结中提取临床判断,构建患者治疗轨迹偏好
- 新奖励函数使器官支持恢复天数增加3.2天,休克缓解更快
- 适合做动态医疗决策的强化学习研究者使用
在医疗强化学习中设计奖励函数面临挑战,因临床结果稀疏、延迟且难以明确设定。结构化临床数据虽能反映生理状态,却常忽略治疗反应、恢复过程与干预负担等纵向信息。相比之下,临床病历包含医生对疾病进展、治疗效果和康复的长期评估,可作为超越预定义指标的轨迹级监督信号。本文提出临床病历引导的偏好奖励(CN-PR)框架,利用大语言模型从出院小结中生成轨迹质量评分,并通过成对偏好构建训练信号,以偏好优化方式学习奖励函数。为应对病历信息量差异,引入任务相关性信号对监督权重进行调节。在离线强化学习场景下评估动态脓毒症治疗策略,所学奖励与轨迹质量评分呈强单调一致,生成策略显著改善恢复相关结果:器官支持免用天数平均增加3.2天,休克缓解时间缩短,同时死亡率与基于结果的基准方法相当。外部验证结果一致。表明临床病历是动态治疗方案中可扩展且表达丰富的监督来源。
原文摘要 · Abstract (English)
Designing reward functions for reinforcement learning (RL) in healthcare remains challenging because clinically meaningful outcomes are sparse, delayed, and difficult to explicitly specify. Although structured clinical data capture physiologic states, they often fail to reflect broader aspects of patient trajectories such as treatment response, recovery dynamics, and intervention burden. Clinical narratives, by contrast, encode longitudinal clinician assessments of disease progression, treatment effectiveness, and recovery, providing a potential source of trajectory-level supervision beyond predefined outcome metrics. We propose Clinical Narrative-informed Preference Rewards (CN-PR), a framework that learns reward functions directly from discharge summaries by treating clinical narratives as scalable supervision for trajectory-level preferences. Using a large language model, we derive trajectory quality scores and construct pairwise preferences between patient trajectories to learn rewards through preference-based optimization. To account for variability in narrative informativeness, we incorporate a task relevance signal that weights supervision according to its relevance to the downstream decision-making task. We evaluate CN-PR in dynamic sepsis treatment using offline RL. The learned reward demonstrated strong monotonic alignment with trajectory quality scores and produced policies associated with improved recovery-related outcomes, including increased organ support-free days and faster shock resolution, while maintaining mortality performance comparable to outcome-based reward baselines. These findings were preserved under external validation. Our results suggest that clinical narratives provide a scalable and expressive source of supervision for reward learning in dynamic treatment regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。