arXiv:2505.23756cs.CV2025-05NeurIPS被引 2

用无姿态图像实现室内3D物体定位与建图,精度超越主流方法。

Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping

  • 以3D框为几何基元,从无姿态图像中联合估计相机位姿与物体轨迹。
  • 在CA-1M和ScanNet++上优于基于点云或体素的全局方法,地图质量显著提升。
  • 适合需要稀疏、参数化建图的场景,如机器人导航与智能空间构建。

我们重新审视场景级3D物体检测,将其视为以3D方向框为几何基础的、兼具定位与建图能力的物体中心框架。现有3D检测方法依赖已知相机位姿进行全局处理,而本文提出的Rooms from Motion(RfM)方法则在无姿态图像集合上运行。通过将结构光重建中的标准2D关键点匹配器替换为基于图像生成3D框的物体中心匹配器,RfM同时估计度量相机位姿、物体轨迹,并最终生成全局语义3D物体地图。当存在先验位姿时,可通过优化全局3D框与单个观测的一致性进一步提升地图质量。RfM在CA-1M和ScanNet++数据集上展现出强定位性能,其生成地图质量优于领先的基于点的方法和多视角3D物体检测方法,尽管后者依赖点云或密集体积的过参数化。该方法实现了通用的物体中心表示,不仅扩展了Cubify Anything至完整场景,还支持本质稀疏的定位与与物体数量成比例的参数化建图。

原文摘要 · Abstract (English)

We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence of metric camera poses, our method, Rooms from Motion (RfM) operates on a collection of un-posed images. By replacing the standard 2D keypoint-based matcher of structure-from-motion with an object-centric matcher based on image-derived 3D boxes, we estimate metric camera poses, object tracks, and finally produce a global, semantic 3D object map. When a priori pose is available, we can significantly improve map quality through optimization of global 3D boxes against individual observations. RfM shows strong localization performance and subsequently produces maps of higher quality than leading point-based and multi-view 3D object detection methods on CA-1M and ScanNet++, despite these global methods relying on overparameterization through point clouds or dense volumes. Rooms from Motion achieves a general, object-centric representation which not only extends the work of Cubify Anything to full scenes but also allows for inherently sparse localization and parametric mapping proportional to the number of objects in a scene.

3D检测无姿态重建物体中心稀疏建图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。