用视频生成训练机器人策略,少样本也能高鲁棒性
Video Generators are Robot Policies
- 用视频与动作联合生成框架,端到端训练
- 仅需少量演示数据即可实现高泛化能力
- 适合追求少样本高效训练的机器人研究者
尽管灵巧操作取得显著进展,当前视觉-运动策略仍受两大挑战制约:在感知或行为分布变化下难以泛化,且性能受限于人类示范数据规模。本文提出视频策略(Video Policy),一个将视频生成与动作生成结合的模块化框架,可端到端训练。结果表明,学习生成机器人行为视频能以极小示范数据量提取有效策略,显著提升鲁棒性与样本效率。该方法在仿真和真实世界中均展现出对未见物体、背景及任务的强大泛化能力。进一步发现,任务成功率与生成视频质量密切相关,无需动作标注的视频数据对新任务泛化具有关键作用。借助大规模视频生成模型,本方法性能优于传统行为克隆,为更可扩展、数据高效的机器人策略学习开辟新路径。
原文摘要 · Abstract (English)
Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is constrained by the size of human demonstration data. In this paper, we use video generation as a proxy for robot policy learning to address both limitations simultaneously. We propose Video Policy, a modular framework that combines video and action generation that can be trained end-to-end. Our results demonstrate that learning to generate videos of robot behavior allows for the extraction of policies with minimal demonstration data, significantly improving robustness and sample efficiency. Our method shows strong generalization to unseen objects, backgrounds, and tasks, both in simulation and the real world. We further highlight that task success is closely tied to the generated video, with action-free video data providing critical benefits for generalizing to novel tasks. By leveraging large-scale video generative models, we achieve superior performance compared to traditional behavior cloning, paving the way for more scalable and data-efficient robot policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。