arXiv:2508.15972cs.ROcs.CV2025-08被引 5

无需3D模型,用扩散模型不确定性实现零样本6D姿态估计

UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation

  • 利用预训练扩散模型的3D先验与不确定性指导单视图重建
  • 多视角融合时按不确定性优先级优化,提升姿态精度与重建质量
  • 适合无现成3D模型的机器人抓取等实际场景

估计新物体的6D姿态是机器人领域的基础挑战,传统方法依赖物体CAD模型,但获取成本高。近期方法虽尝试通过基础模型先验从单/多视图图像重建物体,但常需额外训练或产生幻觉几何。为此,我们提出UnPose,一种零样本、无需3D模型的6D姿态估计与重建框架,利用预训练扩散模型的3D先验和像素级认知不确定性。从单视图RGB-D帧出发,UnPose使用多视图扩散模型结合3D高斯点云(3DGS)表示生成初始3D模型,并附带像素级认知不确定性估计。随着新观测到来,基于扩散模型不确定性引导的多视角融合逐步优化3DGS模型,持续提升姿态估计准确性和3D重建质量。为保证全局一致性,扩散模型生成的视图与后续观测在位姿图中联合优化,形成连贯的3DGS场。大量实验表明,UnPose在6D姿态估计精度与3D重建质量上显著优于现有方法,并在真实机器人操作任务中验证了实用性。

原文摘要 · Abstract (English)

Estimating the 6D pose of novel objects is a fundamental yet challenging problem in robotics, often relying on access to object CAD models. However, acquiring such models can be costly and impractical. Recent approaches aim to bypass this requirement by leveraging strong priors from foundation models to reconstruct objects from single or multi-view images, but typically require additional training or produce hallucinated geometry. To this end, we propose UnPose, a novel framework for zero-shot, model-free 6D object pose estimation and reconstruction that exploits 3D priors and uncertainty estimates from a pre-trained diffusion model. Specifically, starting from a single-view RGB-D frame, UnPose uses a multi-view diffusion model to estimate an initial 3D model using 3D Gaussian Splatting (3DGS) representation, along with pixel-wise epistemic uncertainty estimates. As additional observations become available, we incrementally refine the 3DGS model by fusing new views guided by the diffusion model's uncertainty, thereby continuously improving the pose estimation accuracy and 3D reconstruction quality. To ensure global consistency, the diffusion prior-generated views and subsequent observations are further integrated in a pose graph and jointly optimized into a coherent 3DGS field. Extensive experiments demonstrate that UnPose significantly outperforms existing approaches in both 6D pose estimation accuracy and 3D reconstruction quality. We further showcase its practical applicability in real-world robotic manipulation tasks.

6D姿态估计扩散模型3D重建机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。