用单条专家轨迹生成带约束的新演示,让机器人零样本迁移到真实世界。
Constraint-Preserving Data Generation for Visuomotor Policy Learning
- 将机器人动作拆解为自由空间运动和关键点轨迹约束,实现几何感知生成。
- 在16个仿真任务中平均成功率77%,优于基线50%。
- 适合需要跨物体形状和姿态泛化的机器人操控场景。
大规模示范数据推动了机器人操作的关键突破,但收集成本高昂。本文提出约束保持的数据生成方法(CP-Gen),仅需一条专家轨迹即可生成包含新物体几何与位姿的机器人示范。这些生成的示范用于训练闭环视觉-运动策略,可零样本迁移至真实世界,并泛化于物体几何与位姿的变化。类似以往通过位姿变化生成数据的工作,CP-Gen先将专家示范分解为自由空间运动与机器人技能;但不同于以往,它通过将机器人技能定义为相对于任务相关物体的关键点轨迹约束,实现几何感知生成:机器人或抓取物上的关键点必须跟踪相对于任务相关物体定义的参考轨迹。生成新示范时,对每个任务相关物体采样位姿与几何变换,并将其应用于物体及其关联的关键点或关键点轨迹;随后优化机器人关节配置,使关键点跟踪变换后的轨迹,并规划无碰撞路径至首个优化后的关节状态。在16个仿真任务和4个真实任务上进行实验,涵盖多阶段、非抓握及高精度操作,使用CP-Gen训练的策略平均成功率达77%,显著优于最佳基线的50%。
原文摘要 · Abstract (English)
Large-scale demonstration data has powered key breakthroughs in robot manipulation, but collecting that data remains costly and time-consuming. We present Constraint-Preserving Data Generation (CP-Gen), a method that uses a single expert trajectory to generate robot demonstrations containing novel object geometries and poses. These generated demonstrations are used to train closed-loop visuomotor policies that transfer zero-shot to the real world and generalize across variations in object geometries and poses. Similar to prior work using pose variations for data generation, CP-Gen first decomposes expert demonstrations into free-space motions and robot skills. But unlike those works, we achieve geometry-aware data generation by formulating robot skills as keypoint-trajectory constraints: keypoints on the robot or grasped object must track a reference trajectory defined relative to a task-relevant object. To generate a new demonstration, CP-Gen samples pose and geometry transforms for each task-relevant object, then applies these transforms to the object and its associated keypoints or keypoint trajectories. We optimize robot joint configurations so that the keypoints on the robot or grasped object track the transformed keypoint trajectory, and then motion plan a collision-free path to the first optimized joint configuration. Experiments on 16 simulation tasks and four real-world tasks, featuring multi-stage, non-prehensile and tight-tolerance manipulation, show that policies trained using CP-Gen achieve an average success rate of 77%, outperforming the best baseline that achieves an average of 50%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。