用3D运动流模型让机器人跨硬件学会复杂抓取动作。
3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model
- 从人类和机器人数据中学习3D物体运动流,生成动作指导信号。
- 在11万条数据上训练,实现跨机器人、跨场景的零样本泛化。
- 结合GPT-4o评估与优化,实现闭环动作规划,无需特定硬件训练。
机器人操纵长期面临挑战,而人类能轻松完成如将杯子挂到杯架等复杂操作。核心瓶颈在于缺乏统一的大规模数据集。现有数据集多在简单场景中记录不同动作空间下的机器人行为,难以学习通用鲁棒的动作表征。观察发现,理解物体在三维空间中的运动轨迹是关键线索,该线索与执行者无关,适用于人和各类机器人。为此,我们从人类和机器人操纵数据中学习一个3D流世界模型,预测交互物体未来的3D运动路径,用于指导动作规划。具体地,我们通过移动物体自动检测管道构建大规模3D光流数据集ManiFlow-110k;采用基于视频扩散的世界模型,从这些数据中学习操纵物理规律,生成条件于语言指令的3D光流轨迹。进一步提出流引导渲染机制,渲染预测终态并利用GPT-4o评估其是否符合任务描述,赋予机器人闭环规划能力。最后,将预测的3D光流作为约束,通过优化策略生成一段机器人动作序列。大量实验表明,该方法在多种机器人操纵任务中具备强泛化能力,并实现无需硬件特训的可靠跨体感适应。
原文摘要 · Abstract (English)
Manipulation has long been a challenging task for robots, while humans can effortlessly perform complex interactions with objects, such as hanging a cup on the mug rack. A key reason is the lack of a large and uniform dataset for teaching robots manipulation skills. Current robot datasets often record robot action in different action spaces within a simple scene. This hinders the robot to learn a unified and robust action representation for different robots within diverse scenes. Observing how humans understand a manipulation task, we find that understanding how the objects should move in the 3D space is a critical clue for guiding actions. This clue is embodiment-agnostic and suitable for both humans and different robots. Motivated by this, we aim to learn a 3D flow world model from both human and robot manipulation data. This model predicts the future movement of the interacting objects in 3D space, guiding action planning for manipulation. Specifically, we synthesize a large-scale 3D optical flow dataset, named ManiFlow-110k, through a moving object auto-detect pipeline. A video diffusion-based world model then learns manipulation physics from these data, generating 3D optical flow trajectories conditioned on language instructions. With the generated 3D object optical flow, we propose a flow-guided rendering mechanism, which renders the predicted final state and leverages GPT-4o to assess whether the predicted flow aligns with the task description. This equips the robot with a closed-loop planning ability. Finally, we consider the predicted 3D optical flow as constraints for an optimization policy to determine a chunk of robot actions for manipulation. Extensive experiments demonstrate strong generalization across diverse robotic manipulation tasks and reliable cross-embodiment adaptation without hardware-specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。