用渐进式生成方法提升单目深度估计的几何一致性。
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

- 采用自回归方式逐步构建深度图,从粗到细逐层生成。
- 在多个尺度上注入视觉特征,增强结构一致性。
- 适合需要精细几何细节的3D重建与自动驾驶场景。
扩散模型已成为单目深度估计(MDE)的主流范式,但其隐式假设深度可作为全局平滑场通过迭代去噪恢复,未能显式体现场景几何的分块与尺度依赖特性。实际上,几何结构在空间尺度上是逐步呈现的,粗略布局、表面和边界以层次化方式构建。受此启发,我们提出ARDepth,将深度估计建模为结构化的自回归生成过程。不同于全局优化,ARDepth随着空间分辨率提升逐步构建深度表示。为此,我们引入尺度渐进条件(SPC)在每阶段注入多尺度视觉特征,并设计语义感知引导(SAG)提供场景级语义先验,增强全局结构一致性。两者协同使模型既能捕捉精细局部细节,又保持整体几何连贯性。实验表明,该方法在多尺度下均取得优异性能,生成深度预测具有强结构一致性,验证了自回归生成作为几何建模新范式的潜力。
原文摘要 · Abstract (English)
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In practice, geometric structure emerges progressively across spatial scales, where coarse layout, surfaces, and boundaries are constructed in a hierarchical manner. Motivated by this observation, we introduce ARDepth, which formulates depth estimation as structured auto-regressive generation. Instead of recovering depth through global refinement, ARDepth progressively constructs depth representations as spatial resolution increases. To support this generative process, we introduce Scale-Progressive Conditioning (SPC) to inject multi-scale visual features at each generation stage, and Semantic-Aware Guidance (SAG) to provide scene-level semantic priors that enhance global structural consistency. Together, these designs enable the model to capture fine-grained local details while maintaining coherent global geometry. Empirical results demonstrate that our approach achieves strong performance and produces structurally consistent depth predictions across scales, validating auto-regressive generation as a promising alternative paradigm for geometric modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。