arXiv:2603.18965cs.LGstat.ML2026-03

用未来状态动作特征的熵设计奖励,提升探索效率。

Maximum-Entropy Exploration with Future State-Action Visitation Measures

  • 基于未来状态动作特征的熵设计内在奖励。
  • 实验显示单条轨迹特征覆盖更好,收敛更快。
  • 适合需要高效探索的强化学习任务。

最大熵强化学习通过提供与熵成比例的内在奖励来激励智能体探索状态和动作。本文研究一种新的内在奖励机制,该奖励与未来时间步中访问的状态动作特征的折扣分布熵成比例。这一方法基于两个结论:首先,这些内在奖励的期望总和是初始状态出发轨迹中状态动作特征折扣分布熵的下界,与另一种最大熵目标相关;其次,内在奖励定义所用的分布是压缩算子的不动点,因此可离线估计。实验表明,新目标在单条轨迹内提升了特征覆盖率,尽管整体轨迹间特征覆盖率略有下降(符合下界预期),同时显著加快了仅探索型智能体的学习收敛速度。在多数基准测试上,控制性能与其他方法相当。

原文摘要 · Abstract (English)

Maximum entropy reinforcement learning motivates agents to explore states and actions to maximize the entropy of some distribution, typically by providing additional intrinsic rewards proportional to that entropy function. In this paper, we study intrinsic rewards proportional to the entropy of the discounted distribution of state-action features visited during future time steps. This approach is motivated by two results. First, we show that the expected sum of these intrinsic rewards is a lower bound on the entropy of the discounted distribution of state-action features visited in trajectories starting from the initial states, which we relate to an alternative maximum entropy objective. Second, we show that the distribution used in the intrinsic reward definition is the fixed point of a contraction operator and can therefore be estimated off-policy. Experiments highlight that the new objective leads to improved visitation of features within individual trajectories, in exchange for slightly reduced visitation of features in expectation over different trajectories, as suggested by the lower bound. It also leads to improved convergence speed for learning exploration-only agents. Control performance remains similar across most methods on the considered benchmarks.

强化学习探索机制最大熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。