让机器人记住场景空间布局,提升复杂操作能力
mindmap: Spatial Memory in Deep Feature Maps for 3D Action Policies
- 用深度特征图实现3D环境的空间记忆机制
- 在模拟任务中显著优于无记忆的先进方法
- 适合研究具身智能与长期视觉记忆的学者
端到端学习的机器人控制策略(以神经网络结构)已成为机器人操作的有前景方向。许多常见任务要求物体进出机器人的视野范围,此时空间记忆——即对场景空间构成的持久记忆能力——至关重要。然而,如何在机器人学习系统中构建此类机制仍是开放问题。我们提出 mindmap(基于深度特征图的空间记忆3D动作策略),一种3D扩散策略,其通过环境的语义3D重建生成机器人轨迹。仿真实验表明,该方法在状态领先方法难以解决的任务上表现优异。我们开源了重建系统、训练代码和评估任务,以推动该方向的研究。
原文摘要 · Abstract (English)
End-to-end learning of robot control policies, structured as neural networks, has emerged as a promising approach to robotic manipulation. To complete many common tasks, relevant objects are required to pass in and out of a robot's field of view. In these settings, spatial memory - the ability to remember the spatial composition of the scene - is an important competency. However, building such mechanisms into robot learning systems remains an open research problem. We introduce mindmap (Spatial Memory in Deep Feature Maps for 3D Action Policies), a 3D diffusion policy that generates robot trajectories based on a semantic 3D reconstruction of the environment. We show in simulation experiments that our approach is effective at solving tasks where state-of-the-art approaches without memory mechanisms struggle. We release our reconstruction system, training code, and evaluation tasks to spur research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。