arXiv:2603.27455cs.CV2026-03被引 1

无需标注数据,自监督重建3D几何与相机参数

From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis

  • 用自预测相机参数渲染新视角,实现无标注自监督训练
  • 通过深度感知高斯表示和掩码注意力,稳定优化收敛
  • 兼容主流3D重建架构,适合无先验的开放场景重建

本文提出NAS3R,一种无需真实标注或预训练先验的自监督前馈框架,联合学习显式3D几何与相机参数。训练时,NAS3R从未校准、未对齐的上下文视图重建3D高斯,并利用自预测相机参数渲染目标视图,仅依赖2D光度监督实现自监督训练。为确保稳定收敛,该框架将重建与相机预测整合于共享变压器主干中,采用掩码注意力机制调控,并引入基于深度的高斯建模,提升优化稳定性。该方法兼容当前主流监督式3D重建架构,可在有预训练先验或内在信息时集成使用。大量实验表明,NAS3R在自监督方法中表现领先,建立了可扩展且几何感知的无约束数据3D重建范式。代码与模型已公开于https://ranrhuang.github.io/nas3r/。

原文摘要 · Abstract (English)

In this paper, we introduce NAS3R, a self-supervised feed-forward framework that jointly learns explicit 3D geometry and camera parameters with no ground-truth annotations and no pretrained priors. During training, NAS3R reconstructs 3D Gaussians from uncalibrated and unposed context views and renders target views using its self-predicted camera parameters, enabling self-supervised training from 2D photometric supervision. To ensure stable convergence, NAS3R integrates reconstruction and camera prediction within a shared transformer backbone regulated by masked attention, and adopts a depth-based Gaussian formulation that facilitates well-conditioned optimization. The framework is compatible with state-of-the-art supervised 3D reconstruction architectures and can incorporate pretrained priors or intrinsic information when available. Extensive experiments show that NAS3R achieves superior results to other self-supervised methods, establishing a scalable and geometry-aware paradigm for 3D reconstruction from unconstrained data. Code and models are publicly available at https://ranrhuang.github.io/nas3r/.

3D重建自监督高斯泼溅相机估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。