无需标注数据,直接在3D空间重建图像,学习更准确的三维感知表示。
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- 直接在3D空间进行自监督重建,避免间接推断的偏差。
- 在姿态估计任务上超越RayZer,部分任务媲美甚至超过全监督模型。
- 适合需要强3D理解能力的视觉预训练场景,如机器人、自动驾驶。
自监督预训练推动了语言、2D图像和视频领域基础模型的快速发展,但在从多视角图像中学习3D感知表征方面仍不充分。本文提出E-RayZer,一种直接从无标注图像中学习几何结构化表征的自监督3D视觉模型。与先前方法(如RayZer)通过潜在空间视图合成间接推断3D不同,E-RayZer在3D空间中显式执行自监督3D重建,消除捷径解,获得真正具备3D感知能力的表征。为保障收敛与可扩展性,引入细粒度学习课程,按难易程度组织样本,并在无监督条件下融合异构数据源。实验表明,E-RayZer在姿态估计任务上显著优于RayZer,性能达到甚至超越全监督重建模型如VGGT;其学习到的表征在3D下游任务中也优于DINOv3、CroCo v2、VideoMAE V2和RayZer等主流视觉预训练模型,确立了其作为空间视觉预训练新范式的潜力。
原文摘要 · Abstract (English)
Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations from multi-view images. In this paper, we present E-RayZer, a self-supervised 3D vision model that learns geometrically grounded representations directly from unlabeled images. Unlike prior self-supervised methods such as RayZer, which infer 3D indirectly through latent-space view synthesis, E-RayZer operates directly in 3D space, performing self-supervised 3D reconstruction with Explicit geometry. This formulation eliminates shortcut solutions and yields representations that are 3D-aware. To ensure convergence and scalability, we introduce a fine-grained learning curriculum that organizes training from easy to hard samples and harmonizes heterogeneous data sources without any supervision. Experiments show that E-RayZer significantly outperforms RayZer on pose estimation and matches or sometimes surpasses fully supervised reconstruction models such as VGGT. Furthermore, its learned representations outperform leading visual pre-training models (e.g., DINOv3, CroCo v2, VideoMAE V2, and RayZer) on 3D downstream tasks, establishing E-RayZer as a promising paradigm for spatial visual pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。