用统一表示学习提升机器人在复杂环境中的通用任务能力
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- 融合视觉理解与动态预测,通过百万级教学视频预训练
- 在仿真和真实场景中分别提升9%和12%的任务成功率
- 适合需要跨任务泛化的机器人智能体研究
构建能应对开放环境中多样化任务的通用机器人策略是机器人领域的核心挑战。以往方法(VLA)通常基于视觉-语言模型或生成模型构建通用策略,但视觉-语言预训练提供的语义理解与视觉生成预训练提供的视觉动态建模对具身机器人均至关重要。近期统一生成与理解的模型已展示出在大规模预训练下兼具理解和生成的强大能力。我们提出,机器人策略学习也可受益于理解、规划与连续未来表征学习的结合。为此,我们引入UniJEPA,通过在超过100万条互联网规模的教学操作视频上预训练,获得对高维视觉特征的动态建模能力;随后在机器人具身数据上微调,学习从预测表征到动作标记的映射。大量实验表明,该方法在仿真环境和真实世界分布外任务中分别实现9%和12%的性能提升。
原文摘要 · Abstract (English)
Building generalist robot policies that can handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage knowledge from large-scale pretraining, prior work (VLA) has typically built generalist policies either on top of vision-language understanding models (VLMs) or generative models. However, both semantic understanding from vision-language pretraining and visual dynamics modeling from visual-generation pretraining are crucial for embodied robots. Recent unified models of generation and understanding have demonstrated strong capabilities in both comprehension and generation through large-scale pretraining. We posit that robotic policy learning can likewise benefit from the combined strengths of understanding, planning, and continuous future representation learning. Building on this insight, we introduce UniJEPA, which acquires the ability to dynamically model high-dimensional visual features through pretraining on over 1M internet-scale instructional manipulation videos. Subsequently, UniJEPA is fine-tuned on data collected from the robot embodiment, enabling the learning of mappings from predictive representations to action tokens. Extensive experiments show our approach consistently outperforms baseline methods in terms of 9\% and 12\% across simulation environments and real-world out-of-distribution tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。