用视频生成模型模拟机器人动作,低成本评估复杂策略。
Scalable Policy Evaluation with Video World Models
- 用动作条件视频模型构建可扩展的虚拟世界,避免真实测试
- 在多个指标上实现真实策略值与预测值高度相关(相关性>0.8)
- 适合需快速验证多任务策略的研究者,尤其关注仿真泛化
训练通用机器人操作策略展现出巨大潜力,可实现语言驱动、跨场景的多任务行为。然而,评估这些策略仍面临挑战:真实环境测试成本高、耗时长且存在安全风险,还需频繁重置环境。手动构建和填充仿真环境也未能解决此问题,主要因工程投入大及仿真与现实之间在物理和渲染方面存在显著差距。本文探索使用动作条件视频生成模型作为可扩展的世界模型学习方式,用于策略评估。我们展示了如何将动作条件引入现有预训练视频生成模型中,从而利用互联网规模的自然在线视频进行预训练,无需大量成对的视频-动作数据(该数据收集成本高昂)。论文还研究了数据集多样性、预训练权重以及常见失败案例对评估流程的影响。实验表明,在政策排序、实际策略值与预测值的相关性等多项指标上,该方法均表现出色,证明其在无需真实交互的情况下有效评估策略的可行性。
原文摘要 · Abstract (English)
Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult because real-world testing is expensive, time-consuming, and labor-intensive. It also requires frequent environment resets and carries safety risks when deploying unproven policies on physical robots. Manually creating and populating simulation environments with assets for robotic manipulation has not addressed these issues, primarily due to the significant engineering effort required and the substantial sim-to-real gap, both in terms of physics and rendering. In this paper, we explore the use of action-conditional video generation models as a scalable way to learn world models for policy evaluation. We demonstrate how to incorporate action conditioning into existing pre-trained video generation models. This allows leveraging internet-scale in-the-wild online videos during the pre-training stage and alleviates the need for a large dataset of paired video-action data, which is expensive to collect for robotic manipulation. Our paper examines the effect of dataset diversity, pre-trained weights, and common failure cases for the proposed evaluation pipeline. Our experiments demonstrate that across various metrics, including policy ranking and the correlation between actual policy values and predicted policy values, these models offer a promising approach for evaluating policies without requiring real-world interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。