arXiv:2604.00548cs.CV2026-04中稿 · CVPR被引 2

无需昂贵3D标注,用单目深度和图像对应关系训练3D重建模型

Reliev3R: Relieving Feed-forward Reconstruction from Multi-View Geometric Annotations

  • 利用预训练模型零样本预测的单目相对深度和稀疏对应关系获取3D信息
  • 设计感知模糊性的相对深度损失与基于三角几何的重投影损失
  • 仅用少量数据训练即可媲美全监督模型,适合低资源3D重建场景

近年来,前馈重建模型(FFRMs)在重建质量与下游任务适应性方面展现出巨大潜力。然而,其对多视角几何标注(如3D点云和相机位姿)的过度依赖,使全监督训练难以规模化。本文提出Reliev3R,一种从零开始训练FFRMs的弱监督范式,无需代价高昂的多视角几何标注。该方法摆脱对几何传感数据和计算密集型结构光恢复处理的依赖,直接从预训练模型的零样本预测中获取单目相对深度与图像稀疏对应关系作为3D知识源。Reliev3R核心包含一个感知模糊性的相对深度损失和基于三角几何的重投影损失,以促进多视角几何一致性监督。仅使用较少数据训练,Reliev3R便达到全监督模型的性能水平,朝着低成本3D重建监督与可扩展的FFRMs迈出关键一步。

原文摘要 · Abstract (English)

With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However, the excessive reliance on multi-view geometric annotations, e.g. 3D point maps and camera poses, makes the fully-supervised training scheme of FFRMs difficult to scale up. In this paper, we propose Reliev3R, a weakly-supervised paradigm for training FFRMs from scratch without cost-prohibitive multi-view geometric annotations. Relieving the reliance on geometric sensory data and compute-exhaustive structure-from-motion preprocessing, our method draws 3D knowledge directly from monocular relative depths and image sparse correspondences given by zero-shot predictions of pretrained models. At the core of Reliev3R, we design an ambiguity-aware relative depth loss and a trigonometry-based reprojection loss to facilitate supervision for multi-view geometric consistency. Training from scratch with the less data, Reliev3R catches up with its fully-supervised sibling models, taking a step towards low-cost 3D reconstruction supervisions and scalable FFRMs.

3D重建弱监督单目深度几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。