arXiv:2502.10606cs.CVcs.RO2025-02被引 10

无需预存模型,用扩散模型快速生成3D形状并实时估算物体6自由度位姿

HIPPo: Harnessing Image-to-3D Priors for Model-free Zero-shot 6D Pose Estimation

  • 利用扩散模型生成单图对应的3D网格,实现无须预训练的快速建模
  • 通过持续优化几何与外观,将生成的初始网格逐步替换为更精确的在线观测结果
  • 适合机器人在无先验信息时快速响应新物体,特别适用于真实场景中的即时交互

本工作聚焦于机器人应用中的模型无关零样本6自由度物体位姿估计。现有方法虽能精确估计物体位姿,但严重依赖精心准备的CAD模型或参考图像,而这些准备工作耗时且费力。在真实场景中,3D模型或参考图像往往无法提前获取,亟需即时响应。为此,我们提出新型框架HIPPo,通过利用扩散模型中的图像到3D先验,无需依赖定制化的CAD模型或参考图像,实现模型无关的零样本6D位姿估计。具体而言,我们构建了基于多视角扩散模型和3D重建基础模型的HIPPo Dreamer,可仅凭一瞥在数秒内生成任意未见物体的3D网格。随着更多观测数据的获取,我们提出一种测量引导方案,联合优化物体几何与外观,逐步用更可靠的在线观测替代初始扩散先验。由此,HIPPo能够即时估计并追踪新物体的6D位姿,并保持完整网格以供立即机器人应用。在多个基准上的充分实验表明,当先验参考图像有限时,HIPPo优于现有最先进方法。

原文摘要 · Abstract (English)

This work focuses on model-free zero-shot 6D object pose estimation for robotics applications. While existing methods can estimate the precise 6D pose of objects, they heavily rely on curated CAD models or reference images, the preparation of which is a time-consuming and labor-intensive process. Moreover, in real-world scenarios, 3D models or reference images may not be available in advance and instant robot reaction is desired. In this work, we propose a novel framework named HIPPo, which eliminates the need for curated CAD models and reference images by harnessing image-to-3D priors from Diffusion Models, enabling model-free zero-shot 6D pose estimation. Specifically, we construct HIPPo Dreamer, a rapid image-to-mesh model built on a multiview Diffusion Model and a 3D reconstruction foundation model. Our HIPPo Dreamer can generate a 3D mesh of any unseen objects from a single glance in just a few seconds. Then, as more observations are acquired, we propose to continuously refine the diffusion prior mesh model by joint optimization of object geometry and appearance. This is achieved by a measurement-guided scheme that gradually replaces the plausible diffusion priors with more reliable online observations. Consequently, HIPPo can instantly estimate and track the 6D pose of a novel object and maintain a complete mesh for immediate robotic applications. Thorough experiments on various benchmarks show that HIPPo outperforms state-of-the-art methods in 6D object pose estimation when prior reference images are limited.

6D位姿估计扩散模型零样本机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。