arXiv:2409.16283cs.ROcs.CV2024-09被引 178

用人类视频生成让机器人零样本学会新物品和新动作

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

论文配图:Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
图 1 · 摘自论文原文
  • 用预训练模型生成人类操作视频,指导机器人执行
  • 仅需少量真实机器人数据,就能实现对未知物体的操控
  • 无需微调视频模型,适合快速部署到新任务

机器人如何泛化到未见过的物体类型和新动作?本文提出Gen2Act方法,通过网络公开视频数据训练的生成模型,零样本生成人类操作视频,并用单个策略根据生成视频执行任务。相比昂贵的机器人数据收集,该方法仅使用视频预测模型训练数据量的十分之一左右的机器人交互数据进行策略训练,且完全不需微调视频生成模型,直接使用预训练模型生成人类视频。在多种真实场景中验证,Gen2Act成功实现对未见物体的操控和未曾出现的新动作执行。

原文摘要 · Abstract (English)

How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in terms of predicting motion information from web data through human video generation and conditioning a robot policy on the generated video. Instead of attempting to scale robot data collection which is expensive, we show how we can leverage video generation models trained on easily available web data, for enabling generalization. Our approach Gen2Act casts language-conditioned manipulation as zero-shot human video generation followed by execution with a single policy conditioned on the generated video. To train the policy, we use an order of magnitude less robot interaction data compared to what the video prediction model was trained on. Gen2Act doesn't require fine-tuning the video model at all and we directly use a pre-trained model for generating human videos. Our results on diverse real-world scenarios show how Gen2Act enables manipulating unseen object types and performing novel motions for tasks not present in the robot data. Videos are at https://homangab.github.io/gen2act/

机器人操控视频生成零样本泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。