揭示无监督技能学习与目标导向强化学习的统一理论基础。
Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization

- 将两类方法统一为控制最大化框架,阐明其内在联系。
- 证明不同目标导向任务诱导出不兼容的最优策略。
- 给出技能多样性与下游目标敏感性的精确对应关系。
无监督预训练推动了目标导向强化学习(GCRL)的实证进展,但其理论基础仍不清晰。特别是,互信息技能学习(MISL)通过发现行为多样化的技能来支持后续的目标达成任务,但为何这些技能能提升目标达成能力尚不明晰。问题在于,GCRL和MISL均为广义术语:不同任务采用不同的目标达成评估标准,而不同MISL方法优化不同的行为多样性度量。本文通过控制最大化框架统一二者,识别出三种典型的GCRL形式化,并证明它们在相同环境中可能诱导出不兼容的最优策略。然而,三者共享同一解释:表现良好的目标导向策略应使未来轨迹对指令目标高度敏感,具体敏感性定义由任务形式决定。注意到MISL目标可视为类似目标敏感性的技能敏感性度量,我们证明了MISL目标被特定下游目标敏感性所限制。该界限建立了MISL方法与下游GCRL任务之间的精确对应:对每种GCRL形式,都存在一个匹配的MISL目标,使得更丰富的技能带来更高的下游目标敏感性。研究为强化学习预训练提供了理论基础,并具有重要实践意义,例如指导用户根据下游任务类型选择合适的预训练目标。
原文摘要 · Abstract (English)
Unsupervised pretraining has driven empirical advances in goal-conditioned reinforcement learning (GCRL), but its theoretical foundations remain poorly understood. In particular, an influential class of methods, mutual information skill learning (MISL), discovers behaviorally diverse skills that can later be used for downstream goal-reaching. However, it remains a theoretical mystery why skills learned through MISL should support goal-reaching. A subtle challenge is that both GCRL and MISL are umbrella terms: different GCRL tasks use distinct criteria for measuring goal-reaching performance, while different MISL methods optimize distinct notions of behavioral diversity. We address this challenge and unify GCRL and MISL as instances of control maximization. We identify three canonical GCRL formulations and prove that they are fundamentally inequivalent: they can induce incompatible optimal policies even in the same environment. Nevertheless, they all share a common interpretation: a well-performing goal-conditioned policy is one whose future trajectory is highly sensitive to the commanded goal, with the precise notion of sensitivity determined by the GCRL formulation. Noting that MISL objectives can be understood as measures of skill-sensitivity akin to goal-sensitivity, we show that MISL objectives are bounded by formulation-specific downstream goal-sensitivities. These bounds establish a precise correspondence between MISL methods and downstream GCRL tasks: for every GCRL formulation, there exists a matching MISL objective for which more diverse skills afford greater downstream goal sensitivity. Our results thus lay a theoretical foundation for RL pretraining and have important practical implications, such as suggesting which pretraining objectives to use when a user cares about a specific class of downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。