arXiv:2501.02576cs.CV2025-01被引 24

用单步扩散模型提升单目深度估计的精度与速度,兼顾结构与细节。

DepthMaster: Taming Diffusion Models for Monocular Depth Estimation

  • 设计单步扩散模型,通过特征对齐增强语义表征能力。
  • 引入傅里叶增强模块,自适应平衡结构与细节,提升视觉质量。
  • 两阶段训练策略,分别优化全局结构和局部细节,适合高精度深度任务。

单目深度估计在扩散去噪范式下展现出出色的泛化能力,但推理速度较低。现有方法采用单步确定性框架提升效率,却忽视生成与判别特征间的差异,导致性能受限。本文提出DepthMaster,一种专为判别性深度估计设计的单步扩散模型。首先,提出特征对齐模块,融合高质量语义特征以缓解生成特征对纹理细节的过拟合;其次,设计傅里叶增强模块,自适应平衡低频结构与高频细节。采用两阶段训练策略:第一阶段聚焦全局场景结构学习,第二阶段优化细节表现。实验表明,该模型在多个数据集上均达到领先性能,显著优于其他扩散基方法,在泛化性和细节保留方面表现突出。

原文摘要 · Abstract (English)

Monocular depth estimation within the diffusion-denoising paradigm demonstrates impressive generalization ability but suffers from low inference speed. Recent methods adopt a single-step deterministic paradigm to improve inference efficiency while maintaining comparable performance. However, they overlook the gap between generative and discriminative features, leading to suboptimal results. In this work, we propose DepthMaster, a single-step diffusion model designed to adapt generative features for the discriminative depth estimation task. First, to mitigate overfitting to texture details introduced by generative features, we propose a Feature Alignment module, which incorporates high-quality semantic features to enhance the denoising network's representation capability. Second, to address the lack of fine-grained details in the single-step deterministic framework, we propose a Fourier Enhancement module to adaptively balance low-frequency structure and high-frequency details. We adopt a two-stage training strategy to fully leverage the potential of the two modules. In the first stage, we focus on learning the global scene structure with the Feature Alignment module, while in the second stage, we exploit the Fourier Enhancement module to improve the visual quality. Through these efforts, our model achieves state-of-the-art performance in terms of generalization and detail preservation, outperforming other diffusion-based methods across various datasets. Our project page can be found at https://indu1ge.github.io/DepthMaster_page.

深度估计扩散模型单目图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。