用未来状态与动作访问分布的相对熵做探索奖励,提升策略覆盖和控制性能。
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
- 以未来状态和动作访问分布的相对熵作为内在奖励
- 所学策略在状态-动作空间中具有良好覆盖且控制性能优异
- 可离线学习固定点分布,适合复杂任务的高效探索
最大熵强化学习通过提供与分布熵成比例的内在奖励,将探索融入策略学习。本文提出一种新方法:内在奖励函数为未来时间步中状态与动作(或其派生特征)访问分布的相对熵。该方法基于两个结果:第一,最大化内在奖励期望值等价于最大化决策过程状态-动作值函数的下界;第二,内在奖励定义中的分布是压缩算子的不动点。因此,现有算法可被调整以离线学习该不动点并计算内在奖励。我们最终提出一个优化新目标的算法,结果表明所得策略具备良好的状态-动作空间覆盖性,并实现高性能控制。
原文摘要 · Abstract (English)
Maximum entropy reinforcement learning integrates exploration into policy learning by providing additional intrinsic rewards proportional to the entropy of some distribution. In this paper, we propose a novel approach in which the intrinsic reward function is the relative entropy of the discounted distribution of states and actions (or features derived from these states and actions) visited during future time steps. This approach is motivated by two results. First, a policy maximizing the expected discounted sum of intrinsic rewards also maximizes a lower bound on the state-action value function of the decision process. Second, the distribution used in the intrinsic reward definition is the fixed point of a contraction operator. Existing algorithms can therefore be adapted to learn this fixed point off-policy and to compute the intrinsic rewards. We finally introduce an algorithm maximizing our new objective, and we show that resulting policies have good state-action space coverage and achieve high-performance control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。