POGS让机器人持续追踪不规则物体,无需重扫或建模。
Persistent Object Gaussian Splat (POGS) for Tracking Human and Robot Manipulation of Irregularly Shaped Objects
- 用自监督视觉特征和分组信息构建可更新的物体表示
- 实现12次连续物体重置,80%工具扰动后仍能恢复
- 适合需动态抓取、重定位的工业机器人场景
在制造、装配和物流等场景中,追踪并操作动态环境中不规则形状且未曾见过的物体至关重要。近期提出的高斯点云能高效建模物体几何结构,但缺乏面向任务操作的持续状态估计能力。本文提出持久化物体高斯点云(POGS),将语义、自监督视觉特征与物体分组特征嵌入紧凑表示中,支持对扫描物体姿态的持续更新。该系统无需昂贵的重新扫描或物体预先的CAD模型,在完成初始多视角场景捕获与训练后,仅需单个立体相机,结合深度估计与自监督视觉编码器特征,实现物体位姿估计。POGS支持抓取、重定向及自然语言驱动的操作,通过精修位姿估计,实现人引起的物体扰动下序列重置操作以及工具伺服功能,机器人可在工具扰动达30°时仍恢复工具位姿。实验表明,POGS可实现最多12次连续成功物体重置,并在80%的抓握内工具扰动情况下完成恢复。
原文摘要 · Abstract (English)
Tracking and manipulating irregularly-shaped, previously unseen objects in dynamic environments is important for robotic applications in manufacturing, assembly, and logistics. Recently introduced Gaussian Splats efficiently model object geometry, but lack persistent state estimation for task-oriented manipulation. We present Persistent Object Gaussian Splat (POGS), a system that embeds semantics, self-supervised visual features, and object grouping features into a compact representation that can be continuously updated to estimate the pose of scanned objects. POGS updates object states without requiring expensive rescanning or prior CAD models of objects. After an initial multi-view scene capture and training phase, POGS uses a single stereo camera to integrate depth estimates along with self-supervised vision encoder features for object pose estimation. POGS supports grasping, reorientation, and natural language-driven manipulation by refining object pose estimates, facilitating sequential object reset operations with human-induced object perturbations and tool servoing, where robots recover tool pose despite tool perturbations of up to 30°. POGS achieves up to 12 consecutive successful object resets and recovers from 80% of in-grasp tool perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。