解析目标导向强化学习为何有效,揭示其与双重控制的内在联系。
Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
- 基于最优控制理论,分析目标导向奖励的优化差距
- 在不确定环境中,目标导向策略显著优于传统密集奖励
- 适合处理状态不完全可观测的复杂控制问题
目标条件强化学习旨在训练智能体最大化到达目标状态的概率。本文基于最优控制理论对目标条件设置进行分析,推导出经典二次目标与目标条件奖励之间的最优性差距,解释了目标导向强化学习的成功原因以及传统密集奖励为何会失效。随后,在部分可观测马尔可夫决策过程设定下,将状态估计与概率奖励相结合,表明目标条件奖励特别适用于双重控制问题。通过强化学习和预测控制技术,在非线性和不确定性环境中验证了目标导向策略的优势。
原文摘要 · Abstract (English)
Goal-conditioned reinforcement learning (RL) concerns the problem of training an agent to maximize the probability of reaching target goal states. This paper presents an analysis of the goal-conditioned setting based on optimal control. In particular, we derive an optimality gap between more classical, often quadratic, objectives and the goal-conditioned reward, elucidating the success of goal-conditioned RL and why classical ``dense'' rewards can falter. We then consider the partially observed Markov decision setting and connect state estimation to our probabilistic reward, making the goal-conditioned reward well suited to dual control problems. The advantages of goal-conditioned policies are validated on nonlinear and uncertain environments using both RL and predictive control techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。