arXiv:2511.10647cs.CV2025-11被引 590

用任意视角图恢复精准三维空间,效果超越前代模型。

Depth Anything 3: Recovering the Visual Space from Any Views

  • 仅用一个普通Transformer作为主干,不依赖特殊结构。
  • 在多视角几何任务中平均提升44.3%姿态精度与25.1%几何精度。
  • 适合需要高精度三维重建的视觉研究者和开发者。

我们提出Depth Anything 3(DA3),一种可从任意数量视觉输入中预测空间一致几何的模型,无论是否已知相机位姿。为实现最小化建模,DA3得出两大关键发现:单一普通Transformer(如基础DINO编码器)即可作为主干而无需架构特化;单一深度-射线预测目标足以替代复杂的多任务学习。通过教师-学生训练范式,模型在细节与泛化能力上达到与Depth Anything 2(DA2)相当水平。我们构建了一个涵盖相机位姿估计、任意视角几何与视觉渲染的新视觉几何基准。在该基准上,DA3在所有任务中均达到新SOTA,相比先前最优模型VGGT,平均提升44.3%的姿态准确率与25.1%的几何准确率。此外,在单目深度估计上亦优于DA2。所有模型均仅使用公开学术数据集训练。

原文摘要 · Abstract (English)

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new visual geometry benchmark covering camera pose estimation, any-view geometry and visual rendering. On this benchmark, DA3 sets a new state-of-the-art across all tasks, surpassing prior SOTA VGGT by an average of 44.3% in camera pose accuracy and 25.1% in geometric accuracy. Moreover, it outperforms DA2 in monocular depth estimation. All models are trained exclusively on public academic datasets.

三维重建深度估计多视角Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。