无需CAD模型,仅用10张图就能精准估计物体6D位姿。
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
- 从随机立方体出发,用高斯表示重建物体几何结构。
- 在稀疏视图下实现优于现有方法的6D位姿估计精度。
- 适合无可用3D模型的现实场景,如机器人抓取与增强现实。
单目物体位姿估计是计算机视觉与机器人领域的重要任务,通常依赖于精确的2D-3D对应关系,但这些对应常需昂贵且不易获取的CAD模型。近年来,3D高斯溅射(3DGS)为3D重建提供了新可能,但在输入视图较少时性能下降且易过拟合。为此,本文提出SGPose框架,仅需最少10个输入视图即可完成稀疏视角下的物体位姿估计。该方法从随机立方体初始化开始,不依赖SfM生成的初始几何,而是通过回归图像与重建模型间的密集2D-3D对应关系来实现位姿估计。其成功关键在于几何一致的深度监督与在线合成视图扭曲机制。在典型基准测试中,尤其在遮挡挑战性强的LM-O数据集上,SGPose显著优于现有方法,展现出在真实场景中的巨大应用潜力。
原文摘要 · Abstract (English)
Monocular object pose estimation, as a pivotal task in computer vision and robotics, heavily depends on accurate 2D-3D correspondences, which often demand costly CAD models that may not be readily available. Object 3D reconstruction methods offer an alternative, among which recent advancements in 3D Gaussian Splatting (3DGS) afford a compelling potential. Yet its performance still suffers and tends to overfit with fewer input views. Embracing this challenge, we introduce SGPose, a novel framework for sparse view object pose estimation using Gaussian-based methods. Given as few as ten views, SGPose generates a geometric-aware representation by starting with a random cuboid initialization, eschewing reliance on Structure-from-Motion (SfM) pipeline-derived geometry as required by traditional 3DGS methods. SGPose removes the dependence on CAD models by regressing dense 2D-3D correspondences between images and the reconstructed model from sparse input and random initialization, while the geometric-consistent depth supervision and online synthetic view warping are key to the success. Experiments on typical benchmarks, especially on the Occlusion LM-O dataset, demonstrate that SGPose outperforms existing methods even under sparse view constraints, under-scoring its potential in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。