arXiv:2605.27372cs.CV2026-05

用重力对齐坐标系提升点云处理精度,让3D重建更稳定。

G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing

论文配图:G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
图 1 · 摘自论文原文
  • 将点云从相机坐标转为重力对齐坐标,减少旋转自由度
  • 新模型G3T在真实场景中实现高精度的竖直点云与相机位姿预测
  • 适用于需要稳定三维重建的自动驾驶、机器人领域

当前基于前馈的3D重建方法(如VGGT)在以相机为中心的坐标系中预测像素对齐的点云图。然而,这种坐标系并非总是最优。本文提出在重力对齐的竖直坐标系中预测点云图,利用现实场景中普遍存在的结构线索。与相机坐标系不同,重力对齐坐标系在多视角间共享同一垂直轴,显著降低点云间关联所需的旋转自由度。为此,我们引入重力锚定几何变换器(G3T),在重力对齐的3D数据上微调现有模型。G3T可生成高度准确的重力感知预测,包括竖直点云和相机到重力的位姿。此外,我们提出G3T-Long,一种基于子图的增量式3D重建流程,充分利用竖直坐标系带来的旋转自由度减少,显著提升重建精度。

原文摘要 · Abstract (English)

Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coordinate frame is not always optimal. We propose instead to predict pointmaps in upright, gravity-aligned frames that exploit strong structural cues present in many real-world scenes. Unlike camera-centric frames, gravity-aligned frames share a common vertical axis across viewpoints, reducing the rotational degrees of freedom needed to relate pointmaps to one another. To this end, we introduce the Gravity Grounded Geometry Transformer (G3T), fine-tuned from existing models on gravity-aligned 3D data. G3T produces highly accurate gravity-aware predictions, including upright pointmaps and camera-to-gravity poses. We further introduce G3T-Long, a submap-based incremental 3D reconstruction pipeline that leverages the reduced rotational degrees of freedom afforded by upright frames to achieve significantly improved reconstruction accuracy.

3D重建点云处理坐标系对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。