用生成视频替代真实数据,让机器人学会走路
NIL: No-data Imitation Learning by Leveraging Pre-trained Video Diffusion Models
- 用视频扩散模型生成不同形态的运动视频,作为模仿学习的输入
- 在人形机器人上表现优于依赖真实动作捕捉数据的基线方法
- 适合需要快速适配非人类形态机器人的研究与应用
在多样化且非常规形态(如人形机器人、四足动物和昆虫)上获取物理上合理的运动技能,对角色模拟和机器人技术发展至关重要。传统强化学习方法任务和身体特定,需大量奖励函数设计,泛化能力差;模仿学习虽为替代方案,但严重依赖高质量专家演示,而非常规形态难以获取。视频扩散模型可生成从人类到蚂蚁等各类形态的逼真视频。基于此,我们提出一种无需真实数据的技能获取方法——通过2D生成视频学习3D运动技能,具备向非人类形态泛化的能力。具体而言,利用视觉变换器对视频嵌入进行两两距离计算,结合分段视频帧间的相似性作为引导奖励,指导模仿学习过程。我们在包含独特躯体结构的行走任务中验证了该方法。在人形机器人行走任务中,'无数据模仿学习'(NIL)的表现超越了基于3D动作捕捉数据训练的基线模型。结果表明,利用生成式视频模型可实现多样形态下物理合理技能的学习,有效以数据生成替代数据收集。
原文摘要 · Abstract (English)
Acquiring physically plausible motor skills across diverse and unconventional morphologies-including humanoid robots, quadrupeds, and animals-is essential for advancing character simulation and robotics. Traditional methods, such as reinforcement learning (RL) are task- and body-specific, require extensive reward function engineering, and do not generalize well. Imitation learning offers an alternative but relies heavily on high-quality expert demonstrations, which are difficult to obtain for non-human morphologies. Video diffusion models, on the other hand, are capable of generating realistic videos of various morphologies, from humans to ants. Leveraging this capability, we propose a data-independent approach for skill acquisition that learns 3D motor skills from 2D-generated videos, with generalization capability to unconventional and non-human forms. Specifically, we guide the imitation learning process by leveraging vision transformers for video-based comparisons by calculating pair-wise distance between video embeddings. Along with video-encoding distance, we also use a computed similarity between segmented video frames as a guidance reward. We validate our method on locomotion tasks involving unique body configurations. In humanoid robot locomotion tasks, we demonstrate that 'No-data Imitation Learning' (NIL) outperforms baselines trained on 3D motion-capture data. Our results highlight the potential of leveraging generative video models for physically plausible skill learning with diverse morphologies, effectively replacing data collection with data generation for imitation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。