用物体为中心的神经场实现动作组合生成,让机器人更高效地从示范中学习复杂操作。
Compositional Motion Generation from Demonstration with Object-Centric Neural Fields

- 通过物体级表征连接感知与运动,实现空间-时间联合建模
- 仅需少量数据即可完成长时程仿真任务,性能优于图像基基线
- 适合需要视觉结构理解、泛化能力强的机器人控制场景
组合性通过将复杂行为分解为简单元素的组合,实现可扩展且数据高效的机器人学习。我们提出一种基于示范的生成式学习框架,利用共享的物体级表示将感知与运动相连接,实现机器人行为的组合建模。通过融合规范神经场与隐变量条件变形的物体中心神经表示,对场景进行渲染,能够平滑、一致且可解释地捕捉位置与几何变化。在运动生成方面,采用时间混合专家(MoE)架构,通过门控机制在时间上组合物体条件化的运动原语,生成完整轨迹。这种时空组合性在保持运动原语数据效率的同时,将运动建立在视觉结构基础上,实现了跨多样场景配置的系统性泛化。在仿真中,该模型成功完成长时程操作任务,所需训练数据显著少于其他基于图像的基线方法。真实世界实验进一步验证了方法对噪声的鲁棒性、通过语言分割模型实现类别级泛化的能力,以及直接作用于3D场景表示的潜力。
原文摘要 · Abstract (English)
Compositionality, by organizing complex behavior as combinations of simpler elements, enables robot learning that is scalable and data efficient. Leveraging this principle, we propose a generative learning-from-demonstration framework that enables compositional modeling of robotic behavior by connecting perception and motion through shared object-level representations. We render scenes from object-centric neural representations that integrate canonical neural fields with latent-conditioned deformations, capturing positional and geometric variations in a smooth, consistent, and interpretable way. For motion generation, a temporal mixture-of-experts (MoE) employs a gating mechanism to combine object-conditioned movement primitives over time, producing complete trajectories. This spatial-temporal compositionality maintains the data efficiency of movement primitives while grounding motion in visual structure, enabling systematic generalization across diverse scene configurations. In simulation, long-horizon manipulation tasks are successfully completed using the proposed model, which requires significantly less training data than other image-based baselines. Real-world experiments further demonstrate the method's robustness to noise, its ability to generalize at the category level through language-based segmentation models, and its capacity to operate directly on 3D scene representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。