仅用一张参考图像实现未见物体6自由度位姿估计,支持机器人规模化应用。
Scalable Unseen Objects 6-DoF Absolute Pose Estimation with Robotic Integration
- 基于单张带位姿标签的RGB-D图像,通过状态空间模型迭代对齐点云
- 在6个基准上达到领先性能,真实场景下位姿误差小于1.5°/2.0cm
- 适合工业机器人快速部署,无需3D建模或密集视图采集
姿态估计引导的未见物体6-DoF机器人操作是机器人领域关键任务。然而,现有方法在未见物体上的可扩展性受限,因通常依赖难以获取的CAD模型或密集参考视图。本文提出新任务设定SinRef-6D,仅需机器人操作中捕获的一张带位姿标签的参考RGB-D图像,即可完成未见物体的6-DoF绝对位姿估计。该设定更具可扩展性,但面临大姿态差异及单视图几何信息有限等挑战。为此,我们提出以状态空间模型(SSMs)为骨干的迭代对象空间点对齐策略,通过点与RGB SSM分别捕捉单视图长程空间依赖,具备线性复杂度优势。预训练于合成数据后,SinRef-6D仅凭单参考视图即可估计未见物体位姿。进一步构建软硬件一体化机器人系统并集成该方法,在多个真实场景中验证其有效性。六组基准测试及实际抓取实验表明,该方法在可扩展性方面表现优异,位姿误差低于1.5°/2.0cm。代码与演示视频见https://paperreview99.github.io/SinRef-6DoF-Robotic。
原文摘要 · Abstract (English)
Pose estimation-guided unseen object 6-DoF robotic manipulation is a key task in robotics. However, the scalability of current pose estimation methods to unseen objects remains a fundamental challenge, as they generally rely on CAD models or dense reference views of unseen objects, which are difficult to acquire, ultimately limit their scalability. In this paper, we introduce a novel task setup, referred to as SinRef-6D, which addresses 6-DoF absolute pose estimation for unseen objects using only a single pose-labeled reference RGB-D image captured during robotic manipulation. This setup is more scalable yet technically nontrivial due to large pose discrepancies and the limited geometric and spatial information contained in a single view. To address these issues, our key idea is to iteratively establish point-wise alignment in a common coordinate system with state space models (SSMs) as backbones. Specifically, to handle large pose discrepancies, we introduce an iterative object-space point-wise alignment strategy. Then, Point and RGB SSMs are proposed to capture long-range spatial dependencies from a single view, offering superior spatial modeling capability with linear complexity. Once pre-trained on synthetic data, SinRef-6D can estimate the 6-DoF absolute pose of an unseen object using only a single reference view. With the estimated pose, we further develop a hardware-software robotic system and integrate the proposed SinRef-6D into it in real-world settings. Extensive experiments on six benchmarks and in diverse real-world scenarios demonstrate that our SinRef-6D offers superior scalability. Additional robotic grasping experiments further validate the effectiveness of the developed robotic system. The code and robotic demos are available at https://paperreview99.github.io/SinRef-6DoF-Robotic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。