arXiv:2604.04334cs.LGcs.AI2026-04

BDRL算法提升医疗决策稳定性,让相似患者获得更一致的治疗效果。

Boosted Distributional Reinforcement Learning: Analysis and Healthcare Applications

论文配图:Boosted Distributional Reinforcement Learning: Analysis and Healthcare Applications
图 1 · 摘自论文原文
  • 基于分布强化学习,优化个体治疗效果分布并保证同类患者结果可比。
  • 在高血压管理中,使中高危与脆弱人群的高质量生存年数显著提升。
  • 适合医疗个性化决策、需兼顾公平性与疗效的强化学习场景。

研究人员和从业者正越来越多地利用强化学习优化机器人和医疗等复杂领域的决策。当前工作主要依赖期望值目标,但在高度不确定且存在异质群体的情境下可能不足。尽管分布强化学习能建模结果的完整分布,但可能导致相似代理间的实际收益差异较大。这一问题在医疗中尤为突出,医生需同时管理多个疾病进展和治疗反应各异的患者。本文提出增强型分布强化学习(BDRL),在优化个体结果分布的同时,强制相似代理间结果可比,并分析其收敛性。为稳定学习,引入后更新投影步骤,将其建模为约束凸优化问题,高效将个体结果对齐至指定容差内的高性能参考。我们将该算法应用于美国成人群体的高血压管理,按心血管疾病风险分组,通过模仿各组表现优异的参考个体行为来调整中位及脆弱患者的治疗方案。结果显示,相比基线强化学习方法,BDRL显著提升了高质量生命年的数量与一致性。

原文摘要 · Abstract (English)

Researchers and practitioners are increasingly considering reinforcement learning to optimize decisions in complex domains like robotics and healthcare. To date, these efforts have largely utilized expectation-based learning. However, relying on expectation-focused objectives may be insufficient for making consistent decisions in highly uncertain situations involving multiple heterogeneous groups. While distributional reinforcement learning algorithms have been introduced to model the full distributions of outcomes, they can yield large discrepancies in realized benefits among comparable agents. This challenge is particularly acute in healthcare settings, where physicians (controllers) must manage multiple patients (subordinate agents) with uncertain disease progression and heterogeneous treatment responses. We propose a Boosted Distributional Reinforcement Learning (BDRL) algorithm that optimizes agent-specific outcome distributions while enforcing comparability among similar agents and analyze its convergence. To further stabilize learning, we incorporate a post-update projection step formulated as a constrained convex optimization problem, which efficiently aligns individual outcomes with a high-performing reference within a specified tolerance. We apply our algorithm to manage hypertension in a large subset of the US adult population by categorizing individuals into cardiovascular disease risk groups. Our approach modifies treatment plans for median and vulnerable patients by mimicking the behavior of high-performing references in each risk group. Furthermore, we find that BDRL improves the number and consistency of quality-adjusted life years compared with reinforcement learning baselines.

强化学习医疗决策分布学习个性化治疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。