arXiv:2412.02267cs.CVcs.RO2024-12

无需深度图,用高斯点云+图像对筛选实现单目视频中物体6自由度位姿精准跟踪。

GSGTrack: Gaussian Splatting-Guided Object Pose Tracking from RGB Videos

  • 用3D高斯点云构建可优化的3D表示,联合优化几何与位姿。
  • 在OnePose和HO3D上实现优于现有方法的位姿跟踪精度,重建质量显著提升。
  • 适合机器人抓取等需精确位姿估计的实际场景,尤其在无深度信息时表现突出。

单目RGB视频中未知物体的6自由度位姿跟踪对机器人操作至关重要。然而,现有方法通常依赖精确的深度信息,而真实场景中获取深度信息困难。尽管可使用深度估计算法,但几何误差会导致基于RGBD的位姿跟踪失败。为此,我们提出GSGTrack,一种新型的纯RGB位姿跟踪框架,联合优化几何与位姿。具体地,采用3D高斯点云创建可优化的3D表示,并结合基于图的几何优化来捕捉物体外观特征并精炼几何结构。然而,联合优化易受噪声位姿和几何数据干扰。因此,我们引入物体轮廓损失,缓解像素级损失对位姿噪声的敏感性。为减少不准确深度带来的几何模糊,提出几何一致图像对选择策略,过滤低置信度配对,确保稳健的几何优化。在OnePose和HO3D数据集上的大量实验表明,GSGTrack在6DoF位姿跟踪与物体重建方面均具有显著有效性。

原文摘要 · Abstract (English)

Tracking the 6DoF pose of unknown objects in monocular RGB video sequences is crucial for robotic manipulation. However, existing approaches typically rely on accurate depth information, which is non-trivial to obtain in real-world scenarios. Although depth estimation algorithms can be employed, geometric inaccuracy can lead to failures in RGBD-based pose tracking methods. To address this challenge, we introduce GSGTrack, a novel RGB-based pose tracking framework that jointly optimizes geometry and pose. Specifically, we adopt 3D Gaussian Splatting to create an optimizable 3D representation, which is learned simultaneously with a graph-based geometry optimization to capture the object's appearance features and refine its geometry. However, the joint optimization process is susceptible to perturbations from noisy pose and geometry data. Thus, we propose an object silhouette loss to address the issue of pixel-wise loss being overly sensitive to pose noise during tracking. To mitigate the geometric ambiguities caused by inaccurate depth information, we propose a geometry-consistent image pair selection strategy, which filters out low-confidence pairs and ensures robust geometric optimization. Extensive experiments on the OnePose and HO3D datasets demonstrate the effectiveness of GSGTrack in both 6DoF pose tracking and object reconstruction.

位姿跟踪3D高斯单目视觉机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。