arXiv:2603.16362cs.CVcs.AI2026-03AAAI

40倍提速且保真度高,遥感单目深度估计新框架

$D^3$-RSMDE: 40$\times$ Faster and High-Fidelity Remote Sensing Monocular Depth Estimation

论文配图:$D^3$-RSMDE: 40$\times$ Faster and High-Fidelity Remote Sensing Monocular Depth Estimation
图 1 · 摘自论文原文
  • 用ViT快速生成深度结构先验,替代扩散模型耗时的初始化阶段
  • 轻量U-Net在潜空间迭代细化细节,仅数次迭代即完成优化
  • 相比顶尖模型提速超40倍,显存占用与轻量ViT相当

实时、高保真地从遥感图像中进行单目深度估计对众多应用至关重要,但现有方法在精度与效率间存在显著权衡。尽管基于视觉变换器(ViT)的密集预测方法速度快,但感知质量较差;而扩散模型虽保真度高,计算成本却极高。为此,我们提出适用于遥感单目深度估计的深度细节扩散框架($D^3$-RSMDE),旨在实现速度与质量的最佳平衡。该框架首先利用基于ViT的模块快速生成高质量初步深度图,作为结构先验,有效取代扩散模型中耗时的初始结构生成阶段。在此先验基础上,提出渐进线性混合精化策略(PLBR),使用轻量级U-Net在仅少数迭代内对细节进行优化。整个精化步骤在由变分自编码器(VAE)支持的紧凑潜空间中高效运行。大量实验表明,$D^3$-RSMDE在领先模型Marigold的基础上,使学习感知图像块相似性(LPIPS)指标降低11.85%,同时推理速度提升超过40倍,并保持与轻量级ViT模型相当的显存占用。

原文摘要 · Abstract (English)

Real-time, high-fidelity monocular depth estimation from remote sensing imagery is crucial for numerous applications, yet existing methods face a stark trade-off between accuracy and efficiency. Although using Vision Transformer (ViT) backbones for dense prediction is fast, they often exhibit poor perceptual quality. Conversely, diffusion models offer high fidelity but at a prohibitive computational cost. To overcome these limitations, we propose Depth Detail Diffusion for Remote Sensing Monocular Depth Estimation ($D^3$-RSMDE), an efficient framework designed to achieve an optimal balance between speed and quality. Our framework first leverages a ViT-based module to rapidly generate a high-quality preliminary depth map construction, which serves as a structural prior, effectively replacing the time-consuming initial structure generation stage of diffusion models. Based on this prior, we propose a Progressive Linear Blending Refinement (PLBR) strategy, which uses a lightweight U-Net to refine the details in only a few iterations. The entire refinement step operates efficiently in a compact latent space supported by a Variational Autoencoder (VAE). Extensive experiments demonstrate that $D^3$-RSMDE achieves a notable 11.85% reduction in the Learned Perceptual Image Patch Similarity (LPIPS) perceptual metric over leading models like Marigold, while also achieving over a 40x speedup in inference and maintaining VRAM usage comparable to lightweight ViT models.

深度估计遥感影像扩散模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。