arXiv:2506.09276cs.LGcs.AI2025-06

不依赖奖励和动作,从轨迹学环境结构的最小动作距离

Learning The Minimum Action Distance

  • 从状态轨迹自监督学习最小动作距离(MAD)
  • 在多种环境中准确学习到真实MAD值,优于现有方法
  • 适合需要密集奖励或目标导向强化学习的研究者

本文提出一种仅需状态轨迹即可学习的马尔可夫决策过程(MDP)状态表示框架,无需奖励信号或智能体执行的动作。我们提出学习最小动作距离(MAD),即状态间转换所需的最少动作数,作为刻画环境底层结构的基本度量。MAD自然支持目标条件强化学习和奖励塑造等下游任务,提供密集且几何意义明确的进展度量。所提自监督方法构建嵌入空间,使嵌入状态对之间的距离对应其MAD,兼容对称与非对称近似。我们在涵盖确定性与随机动态、离散与连续状态空间及带噪声观测的多种环境上评估该框架,实验结果表明,该方法在不同设置下均能高效学习精确的MAD表示,且在表示质量上显著优于现有方法。

原文摘要 · Abstract (English)

This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectories, requiring neither reward signals nor the actions executed by the agent. We propose learning the minimum action distance (MAD), defined as the minimum number of actions required to transition between states, as a fundamental metric that captures the underlying structure of an environment. MAD naturally enables critical downstream tasks such as goal-conditioned reinforcement learning and reward shaping by providing a dense, geometrically meaningful measure of progress. Our self-supervised learning approach constructs an embedding space where the distances between embedded state pairs correspond to their MAD, accommodating both symmetric and asymmetric approximations. We evaluate the framework on a comprehensive suite of environments with known MAD values, encompassing both deterministic and stochastic dynamics, as well as discrete and continuous state spaces, and environments with noisy observations. Empirical results demonstrate that the proposed approach not only efficiently learns accurate MAD representations across these diverse settings but also significantly outperforms existing state representation methods in terms of representation quality.

强化学习状态表示自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。