arXiv:2502.01666cs.CVcs.LG2025-02被引 1

用视觉语义增强深度估计,提升复杂场景下精度

Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding

  • 直接从图像特征提取语义信息,替代文本嵌入
  • 在KITTI和Waymo数据集上达到顶尖水平
  • 适合自动驾驶等复杂户外场景的深度感知

单目深度估计旨在从单张RGB图像中预测深度,对自动驾驶、机器人导航、三维重建等应用至关重要。近年来基于学习的方法显著提升了性能。生成模型尤其是Stable Diffusion在大规模多样数据集训练下展现出恢复细节和补全缺失区域的强大能力。然而,依赖文本嵌入的CLIP等模型在复杂户外环境中因缺乏丰富上下文信息而表现受限。为此,本文提出一种基于图像的语义嵌入方法,直接从视觉特征中提取上下文信息,显著提升复杂环境下的深度预测效果。在KITTI和Waymo数据集上的评估表明,该方法性能接近当前最优模型,同时克服了CLIP嵌入在室外场景中的不足。通过直接利用视觉语义,本方法在深度估计任务中表现出更强的鲁棒性和适应性,展示了其在其他视觉感知任务中的应用潜力。

原文摘要 · Abstract (English)

Monocular depth estimation involves predicting depth from a single RGB image and plays a crucial role in applications such as autonomous driving, robotic navigation, 3D reconstruction, etc. Recent advancements in learning-based methods have significantly improved depth estimation performance. Generative models, particularly Stable Diffusion, have shown remarkable potential in recovering fine details and reconstructing missing regions through large-scale training on diverse datasets. However, models like CLIP, which rely on textual embeddings, face limitations in complex outdoor environments where rich context information is needed. These limitations reduce their effectiveness in such challenging scenarios. Here, we propose a novel image-based semantic embedding that extracts contextual information directly from visual features, significantly improving depth prediction in complex environments. Evaluated on the KITTI and Waymo datasets, our method achieves performance comparable to state-of-the-art models while addressing the shortcomings of CLIP embeddings in handling outdoor scenes. By leveraging visual semantics directly, our method demonstrates enhanced robustness and adaptability in depth estimation tasks, showcasing its potential for application to other visual perception tasks.

深度估计视觉语义Stable Diffusion

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。