用拉普拉斯表示法提升决策时规划的长程预测能力
Laplacian Representations for Decision-Time Planning
- 用拉普拉斯表示捕捉多时间尺度的状态距离
- 在OGBench上超越主流基线,尤其擅长长程任务
- 适合需要长视距规划的强化学习场景
基于模型的强化学习中,决策时规划仍面临挑战。状态表示需支持局部代价计算并保持长程结构。本文表明,拉普拉斯表示能有效构建规划用的隐空间,捕捉多时间尺度的状态间距。该表示保留有意义的距离信息,并自然将长程问题分解为子目标,同时缓解长预测时域中的误差累积。基于此特性,我们提出ALPS——一种分层规划算法,在OGBench基准的若干离线目标条件强化学习任务上表现优于常用基线,此前这些任务主要由无模型方法主导。
原文摘要 · Abstract (English)
Planning with a learned model remains a key challenge in model-based reinforcement learning (RL). In decision-time planning, state representations are critical as they must support local cost computation while preserving long-horizon structure. In this paper, we show that the Laplacian representation provides an effective latent space for planning by capturing state-space distances at multiple time scales. This representation preserves meaningful distances and naturally decomposes long-horizon problems into subgoals, also mitigating the compounding errors that arise over long prediction horizons. Building on these properties, we introduce ALPS, a hierarchical planning algorithm, and demonstrate that it outperforms commonly used baselines on a selection of offline goal-conditioned RL tasks from OGBench, a benchmark previously dominated by model-free methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。