arXiv:2605.09364cs.LG2026-05

多尺度预测提升目标导向强化学习的表示鲁棒性

Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning

论文配图:Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning
图 1 · 摘自论文原文
  • 通过多尺度预测监督,让表示同时捕捉局部动态和长程目标结构
  • 在视觉与状态任务上均实现更优表示质量与性能表现
  • 对噪声数据和复杂轨迹拼接场景具有强鲁棒性,适合真实环境应用

本文研究离线目标导向强化学习中的鲁棒表示学习问题。在稀疏奖励场景下,学习与目标潜在变量对齐的表示常导致表示发散,即编码器漂向低维、无关目标的子空间,进而破坏策略学习。为此,我们提出多尺度预测表示框架Ms.PR,强调智能体需在多个尺度(从局部物理动态到长程目标结构)理解环境。该框架利用多尺度预测监督,强制潜空间中状态与目标的一致性。实验表明,Ms.PR显著提升表示质量,在视觉与状态任务上均取得优异性能。此外,该方法在多种现实挑战性数据设置下表现稳健,包括不同任务、轨迹拼接场景及极端噪声条件,持续保持领先水平。

原文摘要 · Abstract (English)

This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations that align state and goal latents is a challenge that frequently culminates in representation divergence where the encoder drifts toward a low-dimensional, goal-agnostic subspace that destabilizes policy learning. We address this issue by showing that an agent must acquire a fundamental understanding of its environment across multiple scales, from local physical dynamics to long-horizon goal-directed structure. Building on this insight, we propose Ms.PR, a framework that leverages multi-scale predictive supervision to enforce goal-directed alignment within the latent space. We demonstrate that Ms.PR leads to improved representation quality and strong performance on both vision and state-based tasks. Furthermore, we show that our approach is exceptionally resilient under realistic, challenging data regimes, maintaining state-of-the-art performance across a wide variety of tasks, trajectory stitching scenarios, and extreme noise conditions.

强化学习表示学习多尺度离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。