arXiv:2606.14389cs.CV2026-06

单张图像中同时重建3D物体并估计多个实例的6D位姿

MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances

  • 基于高斯点云反向渲染,从单视角还原多实例3D结构
  • 在真实和合成场景下实现未见物体的精确重建与位姿估计
  • 适合工业抓取、机器人视觉等需要多实例定位的场景

从单张单目图像中同时进行3D重建与6D物体位姿估计是典型的病态问题。但在工业场景中,同一类物体常随机堆叠于料箱内,单张图像中隐含多个视角的几何信息。我们证明可利用这种隐式多视角结构,同时重建物体3D形态并估计每个可见实例的6D位姿。提出MooMIns方法,基于高斯点云反向建模:不从多相机渲染场景,而是从单相机渲染多个物体实例。模型以SAM3实例分割掩码和改进的运动恢复结构(SfM)流程初始化。相比依赖训练数据先验的深度学习方法,本方法基于图像证据进行真实几何重建,避免了幻觉。在合成与真实料箱抓取场景中评估,验证了对未见过物体的准确重建及个体实例可靠位姿估计能力。

原文摘要 · Abstract (English)

Simultaneous 3D reconstruction and 6D object pose estimation from a single monocular image is an inherently ill-posed problem. In industrial settings, however, multiple instances of an object are often randomly arranged in bins, implicitly providing several views of the same object within a single image. We show that this implicit multi-view geometry can be exploited to simultaneously reconstruct the object in 3D and estimate the 6D pose of each visible object instance. We present MooMIns, a new Gaussian-splatting-based approach that inverts the original Gaussian splatting formulation: instead of rendering a single scene from multiple cameras, we render multiple object instances from a single camera. Our method is initialized with SAM3 instance segmentation masks and a modified Structure from Motion (SfM) pipeline. In contrast to learned monocular depth estimation, we perform true geometry-based reconstruction from image evidence, avoiding hallucinations caused by training data priors. We evaluate MooMIns on synthetic and real bin-picking scenarios, and demonstrate accurate reconstruction of previously unseen objects as well as reliable pose estimation of individual instance

3D重建位姿估计单目视觉工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。