用扩散蒸馏提升单目深度图的精度与细节
SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation
- 用扩散蒸馏融合判别式与生成式方法优点
- 在多个基准上零样本测试均达高精度与清晰边界
- 适合需要真实度与细节兼备的三维场景应用
我们提出 SharpDepth,一种新型单目度量深度估计方法,结合判别式方法(如 Metric3D、UniDepth)的度量准确性与生成式方法(如 Marigold、Lotus)的精细边界锐度。传统判别式模型在稀疏真实深度监督下训练,能准确预测度量深度,但常产生过度平滑或细节不足的深度图。生成式模型则在合成数据与密集真值上训练,生成边界锐利的深度图,但仅提供相对深度且精度较低。我们的方法通过整合度量准确性与细节保留能力,使深度预测兼具度量精确性与视觉清晰性。在标准深度估计基准上的广泛零样本评估验证了 SharpDepth 的有效性,其既能实现高深度精度,又能保持丰富细节,适用于多样真实环境中的高质量深度感知任务。
原文摘要 · Abstract (English)
We propose SharpDepth, a novel approach to monocular metric depth estimation that combines the metric accuracy of discriminative depth estimation methods (e.g., Metric3D, UniDepth) with the fine-grained boundary sharpness typically achieved by generative methods (e.g., Marigold, Lotus). Traditional discriminative models trained on real-world data with sparse ground-truth depth can accurately predict metric depth but often produce over-smoothed or low-detail depth maps. Generative models, in contrast, are trained on synthetic data with dense ground truth, generating depth maps with sharp boundaries yet only providing relative depth with low accuracy. Our approach bridges these limitations by integrating metric accuracy with detailed boundary preservation, resulting in depth predictions that are both metrically precise and visually sharp. Our extensive zero-shot evaluations on standard depth estimation benchmarks confirm SharpDepth effectiveness, showing its ability to achieve both high depth accuracy and detailed representation, making it well-suited for applications requiring high-quality depth perception across diverse, real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。