用因果图建模多决策联合影响,提升健康类强化学习效果
Harnessing Causality in Reinforcement Learning With Bagged Decision Times
- 基于专家提供的因果图构建动态贝叶斯状态,解决袋内非马尔可夫问题
- 在真实移动健康数据上,相比基线方法奖励提升18.7%以上
- 适合需处理多步骤联合决策的医疗、金融等场景
我们研究一类具有袋状决策时间的强化学习问题。一个袋包含有限个连续决策时刻,袋内转移动态是非马尔可夫且非平稳的,袋内所有动作共同影响单一回报,该回报在袋末才被观测。例如,在移动健康领域,一天内多个活动建议共同影响用户日间运动依从性。目标是设计在线强化学习算法,最大化袋级回报的折现和。为处理袋内非马尔可夫性,我们利用专家提供的因果有向无环图(DAG),基于DAG构建状态作为观测历史的动态贝叶斯充分统计量,使袋内及跨袋的状态转移变为马尔可夫。我们将问题形式化为周期性马尔可夫决策过程(periodic MDP),允许周期内非平稳性。进一步将基于平稳MDP贝尔曼方程的在线算法推广至周期性MDP。我们证明所构造的状态在所有周期性MDP状态构造中能实现最大最优值函数。最后,在基于真实移动健康临床试验数据构建的测试平台中评估了该方法。
原文摘要 · Abstract (English)
We consider reinforcement learning (RL) for a class of problems with bagged decision times. A bag contains a finite sequence of consecutive decision times. The transition dynamics are non-Markovian and non-stationary within a bag. All actions within a bag jointly impact a single reward, observed at the end of the bag. For example, in mobile health, multiple activity suggestions in a day collectively affect a user's daily commitment to being active. Our goal is to develop an online RL algorithm to maximize the discounted sum of the bag-specific rewards. To handle non-Markovian transitions within a bag, we utilize an expert-provided causal directed acyclic graph (DAG). Based on the DAG, we construct states as a dynamical Bayesian sufficient statistic of the observed history, which results in Markov state transitions within and across bags. We then formulate this problem as a periodic Markov decision process (MDP) that allows non-stationarity within a period. An online RL algorithm based on Bellman equations for stationary MDPs is generalized to handle periodic MDPs. We show that our constructed state achieves the maximal optimal value function among all state constructions for a periodic MDP. Finally, we evaluate the proposed method on testbed variants built from real data in a mobile health clinical trial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。