从近优轨迹中学习系统动态,提升离线强化学习决策效果
Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
- 利用专家轨迹的近最优性作为约束条件
- 在合成环境与真实医疗场景中显著提升决策性能
- 可判断迁移是否成功,适合医疗、控制等领域应用
我们研究在离线模型强化学习背景下,从近优专家轨迹中估计转移动态 $T^*$。提出一种基于约束的新方法——逆转移学习,将专家轨迹覆盖有限性视为特征:利用专家近乎最优的事实来辅助 $T^*$ 的估计。将这些约束整合进贝叶斯框架。在合成环境及真实医疗场景(如低血压患者在重症监护室管理)中,不仅显著提升了决策表现,还证明后验分布能有效判断迁移成功的可能性。
原文摘要 · Abstract (English)
We consider the problem of estimating the transition dynamics $T^*$ from near-optimal expert trajectories in the context of offline model-based reinforcement learning. We develop a novel constraint-based method, Inverse Transition Learning, that treats the limited coverage of the expert trajectories as a \emph{feature}: we use the fact that the expert is near-optimal to inform our estimate of $T^*$. We integrate our constraints into a Bayesian approach. Across both synthetic environments and real healthcare scenarios like Intensive Care Unit (ICU) patient management in hypotension, we demonstrate not only significant improvements in decision-making, but that our posterior can inform when transfer will be successful.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。