用视频预测训练视觉动态表征,提升机器人控制学习效率。
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
- 以视频预测为预训练任务,用Transformer-CVAE学习视觉动态表征。
- 在多个机器人视觉控制任务上,显著加速下游策略学习。
- 适合需要高效迁移学习的机器人视觉控制场景。
基于纯视频数据进行强化学习(RL)的预训练是一个有价值但具挑战性的问题。尽管真实世界视频易于获取并蕴含丰富的先验世界知识,但缺乏动作标注以及与下游任务间的领域差异限制了其在RL预训练中的应用。为此,我们提出预训练视觉动态表征(PVDR),以弥合视频与下游任务之间的领域差距,实现高效的策略学习。通过采用视频预测作为预训练任务,我们使用基于Transformer的条件变分自编码器(CVAE)学习视觉动态表征,这些表征捕捉了视频中的视觉动态先验知识。该抽象先验知识可直接适配至下游任务,并通过在线适应与可执行动作对齐。我们在一系列机器人视觉控制任务上进行了实验,验证了PVDR是一种有效的视频预训练方式,能显著促进策略学习。
原文摘要 · Abstract (English)
Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence of action annotations and the common domain gap with downstream tasks hinder utilizing videos for RL pre-training. To address the challenge of pre-training with videos, we propose Pre-trained Visual Dynamics Representations (PVDR) to bridge the domain gap between videos and downstream tasks for efficient policy learning. By adopting video prediction as a pre-training task, we use a Transformer-based Conditional Variational Autoencoder (CVAE) to learn visual dynamics representations. The pre-trained visual dynamics representations capture the visual dynamics prior knowledge in the videos. This abstract prior knowledge can be readily adapted to downstream tasks and aligned with executable actions through online adaptation. We conduct experiments on a series of robotics visual control tasks and verify that PVDR is an effective form for pre-training with videos to promote policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。