用语言指导的分步扩散模型提升机器人3D操作的泛化能力
GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D Manipulation
- 分步生成动作,通过语言指令拆解任务并引导子目标
- 在新任务上比顶尖方法高28.63%成功率,训练任务也提升13.36%
- 结合视觉与空间地图,适合真实机器人多任务泛化场景
机器人理解语言指令并执行多样3D操作的能力对机器人学习至关重要。传统基于模仿学习的方法在已见任务上表现良好,但在未见任务上因变化性而受限。近期方法借助大基础模型辅助理解新任务,缓解此问题,但缺乏任务特定学习过程,难以准确理解3D环境,常导致执行失败。本文提出GravMAD,一种子目标驱动、语言条件化的动作扩散框架,融合模仿学习与基础模型优势。该方法根据语言指令将任务分解为子目标,在训练与推理阶段提供辅助引导。训练时引入子目标关键位姿发现以从示范中识别关键子目标;推理时无示范可用,故利用预训练基础模型推断当前任务的子目标。两阶段均生成基于子目标的GravMaps,相比固定3D位置提供更灵活的3D空间引导。在RLBench上的实证评估显示,GravMAD显著优于现有方法,新任务成功率提升28.63%,训练任务提升13.36%。真实机器人任务评估进一步表明,GravMAD可关联视觉信息并泛化至新任务。结果证明其在3D操作中具备强多任务学习与泛化能力。视频演示见:https://gravmad.github.io。
原文摘要 · Abstract (English)
Robots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to variability. Recent approaches leverage large foundation models to assist in understanding novel tasks, thereby mitigating this issue. However, these methods lack a task-specific learning process, which is essential for an accurate understanding of 3D environments, often leading to execution failures. In this paper, we introduce GravMAD, a sub-goal-driven, language-conditioned action diffusion framework that combines the strengths of imitation learning and foundation models. Our approach breaks tasks into sub-goals based on language instructions, allowing auxiliary guidance during both training and inference. During training, we introduce Sub-goal Keypose Discovery to identify key sub-goals from demonstrations. Inference differs from training, as there are no demonstrations available, so we use pre-trained foundation models to bridge the gap and identify sub-goals for the current task. In both phases, GravMaps are generated from sub-goals, providing GravMAD with more flexible 3D spatial guidance compared to fixed 3D positions. Empirical evaluations on RLBench show that GravMAD significantly outperforms state-of-the-art methods, with a 28.63% improvement on novel tasks and a 13.36% gain on tasks encountered during training. Evaluations on real-world robotic tasks further show that GravMAD can reason about real-world tasks, associate them with relevant visual information, and generalize to novel tasks. These results demonstrate GravMAD's strong multi-task learning and generalization in 3D manipulation. Video demonstrations are available at: https://gravmad.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。