融合观察与隐空间预测,提升视觉连续控制的样本效率。
Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

- 结合多步隐空间自预测与下一步观察预测,学习更鲁棒的表征。
- 在DeepMind Control Suite 28个任务上超越现有方法,尤其在狗和人形机器人任务中提升显著。
- 轻量级适配器设计避免隐空间表征过度约束,适合数据受限的复杂视觉控制场景。
从像素中实现高效的策略学习是强化学习中的长期挑战。近期基于动态的表征学习方法通过在潜在空间(自预测)或观测空间(观测预测)进行辅助预测,显著提升了无模型视觉强化学习的样本效率。然而,当前最先进的方法在训练数据有限的复杂视觉控制任务上仍表现不佳。本文认为仅依赖单一预测目标可能不足:观测预测虽使表征扎根于观测级动态,但未直接正则化隐空间表征在长时程上的可预测性。为此,提出观察接地的自预测表征(OG-SPR),一种无模型视觉强化学习算法,学习兼具隐空间时间可预测性和观测级动态一致性的表征。OG-SPR引入双核心辅助目标:多步隐空间自预测与下一步观测预测。实验表明,直接施加隐空间自预测可能过度约束表征且未必提升性能。因此,OG-SPR引入两个轻量级适配器,使共享表征能受益于时间预测信号,而不必直接满足自预测目标。在DeepMind Control Suite的28个视觉控制任务上,OG-SPR在综合性能上优于现有自预测与观测预测方法,尤其在dog、humanoid等高难度任务中表现突出。
原文摘要 · Abstract (English)
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary prediction performed either in latent space (self-prediction) or observation space (observation prediction). However, state-of-the-art methods from both categories still struggle on challenging visual control tasks when training data is limited. We posit that relying on either predictive objective alone may be insufficient. In contrast, observation prediction grounds learned representations in observation-level dynamics, but does not directly regularize the temporal predictability of latent representations over extended horizons. In this paper, we propose Observation-Grounded Self-Predictive Representations (OG-SPR), a model-free visual RL algorithm for continuous control that learns representations that are both temporally predictive in latent space and grounded in observation-level dynamics. OG-SPR incorporates two core auxiliary objectives: multi-step latent self-prediction and next-observation prediction. We empirically show that directly imposing latent self-prediction on the shared representation may over-constrain it and does not necessarily improve performance. To address this issue, OG-SPR introduces two lightweight adapters for latent self-prediction, allowing the shared representation to benefit from temporally predictive signals without being forced to directly satisfy the self-prediction objective. Experiments on 28 visual control tasks from the DeepMind Control Suite show that OG-SPR improves aggregate performance over state-of-the-art self-predictive and observation-predictive RL methods, with particularly pronounced gains in challenging domains such as dog and humanoid.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。