arXiv:2601.11890cs.LG2026-01被引 1

统一了马尔可夫决策中的探索覆盖框架,实现智能优先探索。

From Relative Entropy to Minimax: A Unified Framework for Coverage in MDPs

  • 提出基于状态-动作占据率的凹函数覆盖目标族,统一多种探索策略。
  • 参数ρ增大时,算法自动聚焦最未探索的状态-动作对,极限下等价于最坏情况覆盖。
  • 梯度显式控制,适合需精准探索覆盖的强化学习任务。

在无奖励马尔可夫决策过程(MDP)中,针对性和主动探索状态-动作对至关重要。不同状态-动作对的重要性或探索难度各异,需显式融入探索策略。为此,我们提出一个定义在状态-动作占据率上的加权参数化凹函数覆盖目标族 $U_ρ$,统一了多种广泛研究的目标,包括基于散度的边际匹配、加权平均覆盖和最坏情况(极小极大)覆盖。$U_ρ$ 的凹性捕捉了过度探索的边际递减效应,其梯度具有简洁闭式表达,可显式引导算法优先探索未充分覆盖的状态-动作对。基于此结构,我们设计了一种基于梯度的算法,主动调控占据分布以实现期望的覆盖模式。此外,我们证明当参数 $ρ$ 增大时,探索策略逐渐聚焦于最不被探索的状态-动作对,在极限下恢复最坏情况覆盖行为。

原文摘要 · Abstract (English)

Targeted and deliberate exploration of state--action pairs is essential in reward-free Markov Decision Problems (MDPs). More precisely, different state-action pairs exhibit different degree of importance or difficulty which must be actively and explicitly built into a controlled exploration strategy. To this end, we propose a weighted and parameterized family of concave coverage objectives, denoted by $U_ρ$, defined directly over state--action occupancy measures. This family unifies several widely studied objectives within a single framework, including divergence-based marginal matching, weighted average coverage, and worst-case (minimax) coverage. While the concavity of $U_ρ$ captures the diminishing return associated with over-exploration, the simple closed form of the gradient of $U_ρ$ enables an explicit control to prioritize under-explored state--action pairs. Leveraging this structure, we develop a gradient-based algorithm that actively steers the induced occupancy toward a desired coverage pattern. Moreover, we show that as $ρ$ increases, the resulting exploration strategy increasingly emphasizes the least-explored state--action pairs, recovering worst-case coverage behavior in the limit.

强化学习探索策略覆盖率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。