arXiv:2605.06470cs.LG2026-05

用击中时间重构马尔可夫过程的定向时序几何,提升长程导航规划性能。

Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies

  • 基于击中时间构建希尔伯特空间位移几何,实现非对称距离建模。
  • 理论证明全局击中时间误差受单步转移误差与瞬态谱半径放大影响。
  • 提出IEL算法,适用于离线强化学习中的多阶段规划,适合长期任务建模。

我们提出一种新的算子理论表示学习框架,从击中时间观测中恢复受控马尔可夫过程的定向时序几何。以往方法常产生对称距离或不满足三角不等式,而本框架在潜在线性闭包条件下,学习到将期望击中时间作为潜在位移线性泛函的希尔伯特空间位移几何,并证明其存在且唯一(至有界线性同构)。对于有限维实现,我们证明全局击中时间误差被单步转移误差放大,放大因子为环境的瞬态谱半径。同时提供涵盖近似、统计复杂度和轨迹标签不匹配的有限样本保证。由此推导出的同构嵌入学习(IEL)是一种无目标的基础策略学习算法,结合了HILP风格的一致性目标与显式击中时间回归,确保学习几何反映真实决策进度。该非对称且可组合结构支持鲁棒的图基多阶段规划,适用于长时程导航。实验表明,IEL在离线迷宫运动数据上显著优于现有基础策略学习方法。

原文摘要 · Abstract (English)

We present a new operator-theoretic representation learning framework for offline reinforcement learning that recovers the directed temporal geometry of a controlled Markov process from hitting time observations. While prior art often produces symmetric distances or fails to satisfy the triangle inequality, our framework learns a Hilbert-space displacement geometry where expected hitting times are realized as linear functionals of latent displacements. We prove that this representation exists under latent linear closure and is uniquely identifiable up to a bounded linear isomorphism. For finite-dimensional implementations, we show that global hitting-time error is bounded by one-step transition error amplified by the environment's transient spectral radius. Furthermore, we provide finite-sample guarantees accounting for approximation, statistical complexity, and trajectory-label mismatch. Derived from this theory, we curate Isomorphic Embedding Learning (IEL) as a new goal-agnostic foundation policy learning algorithm that anchors a HILP-style consistency objective with explicit hitting-time regression to ensure that the learned geometry reflects actual decision-time progress. This asymmetric and compositional structure enables robust graph-based multi-stage planning for long-horizon navigation. Our experiments demonstrate that IEL improves the state of the art of learning foundation policy policies from offline maze locomotion data. Our code can be found on https://github.com/MagnusBoock/IEL

强化学习表示学习多阶段规划击中时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。