通过粗粒度动态不确定性约束,让高层智能体更稳健地选择子目标。
S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning

- 用多步聚合的粗粒度动态代替原始动作,降低高层策略的不确定性。
- 在非平稳长时序环境中,性能超越现有先进HRL方法。
- 适合需要稳定规划的复杂长任务,如机器人导航与控制。
层级强化学习(HRL)旨在将战略规划与基础执行分离,已在解决长时程、复杂任务中取得成功,而传统扁平强化学习难以应对。然而,高层智能体依赖稀疏延迟反馈,其表现高度依赖底层执行能力。本文提出一种基于粗粒度动态的内在动机机制,用于优化子目标选择:通过聚合高层时间尺度上的环境转移,构建粗粒度动态模型,并利用混合密度网络(MDN)估算预测不确定性,以抑制高层策略的波动。实验表明,该方法通过密集的动力学感知内在奖励实现风险规避型子目标选择,在非平稳长时序环境中显著优于当前最优的HRL方法。
原文摘要 · Abstract (English)
Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution. It has been widely successful in solving long-horizon and complex tasks, where flat-RL algorithms have difficulty in learning. However, while the low-level agent in HRL benefits from dense feedback and abundant trial opportunities, the high-level agent receives sparse, delayed feedback from the environment and its performance depends on the low-level execution capability. In this paper, we study whether subgoal selection by the high-level agent can be performed more strategically, by providing it with dynamics-aware intrinsic motivation. Since motivation based on primitive transition dynamics would require broad coverage of the state-action space, we propose to use coarse dynamics, i.e., environment transitions aggregated over multiple steps at the temporal scale at which the high-level agent operates. This approach stabilizes the high-level policy by learning to minimize the predictive uncertainty associated with the coarse dynamics, and provides a guided structure for navigation. We model the predictive uncertainty by evaluating different dispersion metrics as approximated by a Mixture Density Network (MDN). Empirically, we observe that a dense, dynamics-aware intrinsic reward leads to risk-averse subgoal selection, enabling it to outperform state-of-the-art HRL methods in non-stationary long-horizon environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。