用物体空间姿态训练轻量机器人策略,无需人工建模。
PRISM-DP: Spatial Pose-based Observations for Diffusion-Policies via Segmentation, Mesh Generation, and Pose Tracking
- 通过分割、网格生成和姿态追踪,从视觉中提取物体空间姿态
- 在真实与仿真环境均优于图像基策略,接近真值状态表现
- 免手动建模,适合开放场景的可扩展机器人控制
扩散策略通过学习去噪动作空间轨迹来生成机器人运动,通常依赖高维RGB图像作为观测,其中包含大量无关信息,需大型模型提取有效模式。相比之下,使用关键物体的空间姿态等结构化观测,可训练参数更少的紧凑策略。然而,在开放集真实环境中准确获取物体姿态仍具挑战,因6D姿态估计与跟踪方法常依赖预先标记或需人工重建的物体网格。本文提出PRISM-DP,利用分割、网格生成与姿态追踪模型,实现直接从任务相关物体空间姿态训练紧凑扩散策略。关键在于,通过网格生成模型,消除了对人工建模的需求,提升了开放环境中的可扩展性。仿真与真实世界实验表明,PRISM-DP优于基于高维图像的策略,并达到与使用真值状态信息训练策略相当的性能。
原文摘要 · Abstract (English)
Diffusion policies generate robot motions by learning to denoise action-space trajectories conditioned on observations. These observations are commonly streams of RGB images, whose high dimensionality includes substantial task-irrelevant information, requiring large models to extract relevant patterns. In contrast, using structured observations like the spatial poses of key objects enables training more compact policies with fewer parameters. However, obtaining accurate object poses in open-set, real-world environments remains challenging, as 6D pose estimation and tracking methods often depend on markers placed on objects beforehand or pre-scanned object meshes that require manual reconstruction. We propose PRISM-DP, an approach that leverages segmentation, mesh generation, and pose tracking models to enable compact diffusion policy learning directly from the spatial poses of task-relevant objects. Crucially, by using a mesh generation model, PRISM-DP eliminates the need for manual mesh creation, improving scalability in open-set environments. Experiments in simulation and the real world show that PRISM-DP outperforms high-dimensional image-based policies and achieves performance comparable to policies trained with ground-truth state information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。