arXiv:2603.11174cs.CV2026-03被引 2

用几何约束提升3D点云重建精度,让模型更准更完整。

GGPT: Geometry Grounded Point Transformer

  • 基于稀疏视图构建精确相机位姿与部分点云,提供几何引导
  • 在部分几何监督下优化点云生成,恢复纹理缺失区域的细节
  • 仅用ScanNet++训练即可跨架构、跨数据集显著超越现有方法

近期前馈网络通过直接从RGB图像预测稠密点云,在稀疏视角3D重建上取得显著进展。然而,由于缺乏显式多视角约束,常出现几何不一致和细粒度精度不足的问题。本文提出几何引导点变换器(GGPT),在前馈重建中引入可靠的稀疏几何引导。首先,基于密集特征匹配与轻量级几何优化,设计改进的运动结构(SfM)流程,高效估计稀疏输入视图下的准确相机位姿与部分3D点云。在此基础上,提出一种几何引导的3D点变换器,在优化的引导编码下,以显式部分几何监督精炼稠密点云。大量实验表明,该方法为融合几何先验与稠密前馈预测提供了合理机制,生成兼具几何一致性与空间完整性的重建结果,能恢复精细结构并填补无纹理区域空缺。仅在ScanNet++上使用VGGT预测进行训练,GGPT在同域与跨域设置下均显著优于当前最优前馈3D重建模型。

原文摘要 · Abstract (English)

Recent feed-forward networks have achieved remarkable progress in sparse-view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limited fine-grained accuracy due to the absence of explicit multi-view constraints. We introduce the Geometry-Grounded Point Transformer (GGPT), a framework that augments feed-forward reconstruction with reliable sparse geometric guidance. We first propose an improved Structure-from-Motion pipeline based on dense feature matching and lightweight geometric optimisation to efficiently estimate accurate camera poses and partial 3D point clouds from sparse input views. Building on this foundation, we propose a geometry-guided 3D point transformer that refines dense point maps under explicit partial-geometry supervision using an optimised guidance encoding. Extensive experiments demonstrate that our method provides a principled mechanism for integrating geometric priors with dense feed-forward predictions, producing reconstructions that are both geometrically consistent and spatially complete, recovering fine structures and filling gaps in textureless areas. Trained solely on ScanNet++ with VGGT predictions, GGPT generalises across architectures and datasets, substantially outperforming state-of-the-art feed-forward 3D reconstruction models in both in-domain and out-of-domain settings.

3D重建点云生成几何引导前馈网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。