arXiv:2502.19204cs.CV2025-02被引 23

不依赖全局归一化,用跨上下文蒸馏提升单目深度估计性能

Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator

  • 提出跨上下文蒸馏,融合全局与局部深度线索
  • 在多个基准数据集上超越现有最佳方法,精度显著提升
  • 适合需要高鲁棒性单目深度估计的研究者与工程师

近期零样本单目深度估计(MDE)通过归一化深度表示统一深度分布,并利用大规模无标签数据进行伪标签蒸馏,显著提升了泛化能力。然而,现有方法依赖全局深度归一化,对所有深度值同等处理,易放大伪标签噪声,降低蒸馏效果。本文系统分析了伪标签蒸馏中的深度归一化策略,发现在此类蒸馏范式下(如共享上下文蒸馏),归一化并非必需,省略可缓解噪声监督影响。我们进一步提出跨上下文蒸馏,结合全局与局部深度线索以提升伪标签质量;同时引入基于扩散模型的教师模型辅助蒸馏,注入互补深度先验,增强监督多样性与鲁棒性。大量实验表明,该方法在多个基准数据集上均显著优于当前最优方法,定量与定性结果均更优。

原文摘要 · Abstract (English)

Recent advances in zero-shot monocular depth estimation(MDE) have significantly improved generalization by unifying depth distributions through normalized depth representations and by leveraging large-scale unlabeled data via pseudo-label distillation. However, existing methods that rely on global depth normalization treat all depth values equally, which can amplify noise in pseudo-labels and reduce distillation effectiveness. In this paper, we present a systematic analysis of depth normalization strategies in the context of pseudo-label distillation. Our study shows that, under recent distillation paradigms (e.g., shared-context distillation), normalization is not always necessary, as omitting it can help mitigate the impact of noisy supervision. Furthermore, rather than focusing solely on how depth information is represented, we propose Cross-Context Distillation, which integrates both global and local depth cues to enhance pseudo-label quality. We also introduce an assistant-guided distillation strategy that incorporates complementary depth priors from a diffusion-based teacher model, enhancing supervision diversity and robustness. Extensive experiments on benchmark datasets demonstrate that our approach significantly outperforms state-of-the-art methods, both quantitatively and qualitatively.

单目深度知识蒸馏扩散模型深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。