arXiv:2608.14157cs.AIcs.LG2026-08

剔除重复病历文本提升强化学习在医疗决策中的表现

Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine

论文配图:Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine
图 1 · 摘自论文原文
  • 用嵌入空间分解或句级差异操作消除病历文本的时间冗余
  • 在真实ICU数据上,新状态表示使强化学习性能显著优于基线
  • 适合关注临床决策支持与多模态强化学习的研究者

机械通气是关键的生命支持手段,需随患者状况动态调整治疗参数。强化学习(RL)为优化这类序列决策提供了前景,但传统方法仅依赖结构化电子健康记录(EHR),忽略了自由文本病历中的重要临床信息。将纵向病历文本纳入RL状态空间面临挑战:文本存在严重的时间冗余,如复制前文、模板化和重复记录,导致时间局部更新被稀释,状态表示质量下降。为此,我们提出一种冗余感知的多模态状态表示框架,在策略学习前显式去除时间上的重复文本。评估了两种计算高效的时序分解策略:(1)基于局部历史子空间奇异值分解的嵌入空间分解;(2)可解释的句级差异操作,在文本编码前过滤已记录过的句子。基于真实重症监护数据,结果表明,去除时间冗余后构建的状态表示,在多种离策略评估方法(模型驱动回放、拟合Q评估、加权重要性采样、加权双重稳健评估)中均显著优于仅使用结构化数据和原始病历文本的基线。研究发现,显式分离新增临床信息与重复文本,能生成更高质量的状态表示,直接提升临床决策支持的强化学习性能。

原文摘要 · Abstract (English)

Mechanical ventilation is a critical life-support intervention, requiring dynamic adjustments to ventilator settings as a patient's condition evolves. While reinforcement learning (RL) offers a promising framework for optimizing these sequential decisions, standard approaches rely primarily on structured electronic health record (EHR) data, missing crucial clinical context recorded in free-text notes. Integrating longitudinal clinical notes into RL state spaces is challenging because notes are heavily inflated by temporal redundancy, such as copy-forward text, templating, and repetitive documentation, which dilutes time-local updates and degrades state representation quality. To address this, we propose a redundancy-aware multimodal state representation framework that explicitly removes duplicated note text over time before policy learning. We evaluate two computationally efficient temporal decomposition strategies for removing duplicated note text: (1) an embedding-space decomposition using singular value decomposition on local history subspaces, and (2) an interpretable sentence-level diff operation that filters out previously documented sentences before text encoding. Using real-world ICU data, we demonstrate that state representations constructed by stripping temporal note redundancy significantly outperform both structured-only and raw-note baselines across multiple off-policy evaluation methods (Model-Based Rollouts, Fitted Q-Evaluation, Weighted Importance Sampling, and Weighted Doubly Robust Evaluation). Our findings show that explicitly isolating new clinical information from repeated note text yields higher-quality state representations and directly improves RL performance for clinical decision support.

强化学习医疗决策多模态病历分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。