用视频扩散模型+轻量适配器,让新机器人仅用20分钟演示就能学会复杂操作。
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
- 用跨平台75万条多视角轨迹继续预训练视频扩散模型,构建通用先验
- 在未见机器人上仅用20分钟人类示范,即实现跨任务/背景/视角的泛化
- 无需密集标注,通过掩码逆动力学模型精准对齐动作空间,适合快速部署
将通用操控能力扩展到新机器人平台仍具挑战:每种平台通常需要大量同质示范数据,而端到端像素到动作的流程在背景和视角变化下易退化。基于视频驱动机器人控制的进展,我们提出Vidar,由一个可泛化的身体化视频扩散模型作为先验,以及一个掩码逆动力学模型(MIDM)作为适配器组成。我们利用在互联网规模预训练的视频扩散模型,并使用来自三个真实机器人平台的75万条多视角轨迹,在身体化领域进行持续预训练。为此,我们引入统一观测空间,联合编码机器人、相机、任务与场景上下文。MIDM模块无需密集标签即可学习与动作相关的像素掩码,将先验锚定至目标机体的动作空间,同时抑制干扰因素。仅需在未见机器人上进行20分钟的人类示范(典型数据量的1%),Vidar便超越现有最先进基线,且能泛化至未见任务、背景和相机布局。结果表明,‘一个先验,多种机体’是一种可扩展的解决方案:强而低成本的视频先验配合极小量的机器人对齐。
原文摘要 · Abstract (English)
Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action pipelines may degenerate under background and viewpoint shifts. Based on previous advances in video-based robot control, we present Vidar, consisting of an embodied video diffusion model as the generalizable prior and a masked inverse dynamics model (MIDM) as the adapter. We leverage a video diffusion model pre-trained at Internet scale, and further continuously pre-train it for the embodied domain using 750K multi-view trajectories collected from three real-world robot platforms. For this embodied pre-training, we introduce a unified observation space that jointly encodes robot, camera, task, and scene contexts. The MIDM module learns action-relevant pixel masks without dense labels, grounding the prior into the target embodiment's action space while suppressing distractors. With only 20 minutes of human demonstrations on an unseen robot (1% of typical data), Vidar outperforms state-of-the-art baselines and generalizes to unseen tasks, backgrounds, and camera layouts. Our results suggest a scalable recipe for "one prior, many embodiments": strong, inexpensive video priors together with minimal on-robot alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。