POLAR通过悲观策略优化,实现更稳健的动态治疗方案学习。
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
- 基于离线数据估计状态转移并量化不确定性,引入悲观惩罚抑制高不确定动作。
- 在合成数据和MIMIC-III上表现优于现有方法,接近最优治疗策略。
- 首个兼具统计与计算保障的模型化DTR算法,适合医疗等高风险决策场景。
动态治疗方案(DTR)为需随个体轨迹动态调整决策的领域(如医疗、教育)提供了系统框架。现有统计方法常依赖强正性假设且对部分数据覆盖不鲁棒;离线强化学习方法多关注平均训练性能,缺乏统计保证,且需复杂优化。为此,我们提出POLAR——一种新型悲观模型化策略学习算法,用于离线DTR优化。POLAR从离线数据中估计转移动态,为每段历史-动作对量化不确定性,并在奖励函数中加入悲观惩罚以抑制高不确定性动作。不同于仅关注平均性能或仅对理想策略提供保证的方法,POLAR直接最小化最终策略的次优性,并提供理论保障,无需复杂极小极大或约束优化。据我们所知,POLAR是首个同时具备统计与计算保障的模型化DTR方法,包含有限样本下策略次优性的界。在合成数据与MIMIC-III数据集上的实验表明,POLAR显著优于现有方法,生成近最优、历史感知的治疗策略。
原文摘要 · Abstract (English)
Dynamic treatment regimes (DTRs) provide a principled framework for optimizing sequential decision-making in domains where decisions must adapt over time in response to individual trajectories, such as healthcare, education, and digital interventions. However, existing statistical methods often rely on strong positivity assumptions and lack robustness under partial data coverage, while offline reinforcement learning approaches typically focus on average training performance, lack statistical guarantees, and require solving complex optimization problems. To address these challenges, we propose POLAR, a novel pessimistic model-based policy learning algorithm for offline DTR optimization. POLAR estimates the transition dynamics from offline data and quantifies uncertainty for each history-action pair. A pessimistic penalty is then incorporated into the reward function to discourage actions with high uncertainty. Unlike many existing methods that focus on average training performance or provide guarantees only for an oracle policy, POLAR directly targets the suboptimality of the final learned policy and offers theoretical guarantees, without relying on computationally intensive minimax or constrained optimization procedures. To the best of our knowledge, POLAR is the first model-based DTR method to provide both statistical and computational guarantees, including finite-sample bounds on policy suboptimality. Empirical results on both synthetic data and the MIMIC-III dataset demonstrate that POLAR outperforms state-of-the-art methods and yields near-optimal, history-aware treatment strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。