通过建模多尺度时间相关性,提升强化学习预训练的表示能力
From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-training

- 设计时间相关性空间,分尺度学习视频时序特征
- 在多个下游任务中提升样本效率与最终性能
- 适合需要高效预训练的强化学习研究者
无监督大规模数据预训练在提升强化学习(RL)样本效率和性能方面展现出巨大潜力。现有方法利用互联网视频中的无动作视频,通过单步状态转移预测和图像重建来学习表征,但倾向于保留像素空间中大量静态信息,忽略小而关键的动态信息。为保留充分信息,需对视频中每个元素给予同等关注。为此,我们提出时间相关性空间以区分各元素。具体实现上,引入多尺度时间对比学习(MTCL)方法,分别建模多尺度时间相关性。该方法可均衡不同元素的关注度,生成更具信息量的表征,有效支持多种下游策略学习任务。实验表明,该方法在多个下游任务中均显著提升样本效率与渐近性能。
原文摘要 · Abstract (English)
Unsupervised pre-training on large-scale datasets has demonstrated significant potential for improving the sample efficiency and performance of Reinforcement Learning (RL). Given the large-scale action-free internet videos, existing methods utilize single-step transition prediction and image reconstruction to learn representations. However, these methods prefer to preserve large-proportion stationary information in the pixel space, neglecting small but crucial information. To preserve enough information in the representation, it is essential to pay equal attention to each element in videos. Specifically, we propose a temporal correlation space to distinguish each element. For implementation, we introduce the Multi-scale Temporal Contrastive Learning (MTCL) method to model multi-scale temporal correlations separately. This approach can balance the attention of different elements and yield more informative representations, effectively supporting policy learning in various downstream tasks. Experimental results demonstrate that our method improves sample efficiency and asymptotic performance across various downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。