用状态间时间距离构建双目标表示,提升强化学习的导航性能。
Dual Goal Representations
- 通过状态间时间距离建模,实现与原始表征无关的目标表示。
- 在20个任务上显著提升离线目标导向强化学习性能。
- 适合追求鲁棒性与泛化能力的强化学习研究者使用。
本文提出用于目标条件强化学习(GCRL)的双目标表示。该表示通过“当前状态到所有其他状态的时间距离集合”来刻画状态,即以状态间关系为依据进行编码。该表示具备多项理论优势:仅依赖环境内在动态,对原始状态表征不变;且可证明包含恢复最优目标到达策略的充分信息,同时能过滤外部噪声。基于此概念,我们设计了一种可与任意现有GCRL算法结合的实用目标表示学习方法。在OGBench任务套件上的多样化实验表明,双目标表示在20个基于状态和像素的任务中均持续提升离线目标到达性能。
原文摘要 · Abstract (English)
In this work, we introduce dual goal representations for goal-conditioned reinforcement learning (GCRL). A dual goal representation characterizes a state by "the set of temporal distances from all other states"; in other words, it encodes a state through its relations to every other state, measured by temporal distance. This representation provides several appealing theoretical properties. First, it depends only on the intrinsic dynamics of the environment and is invariant to the original state representation. Second, it contains provably sufficient information to recover an optimal goal-reaching policy, while being able to filter out exogenous noise. Based on this concept, we develop a practical goal representation learning method that can be combined with any existing GCRL algorithm. Through diverse experiments on the OGBench task suite, we empirically show that dual goal representations consistently improve offline goal-reaching performance across 20 state- and pixel-based tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。