arXiv:2501.13796cs.CV2025-01被引 1

用可学习提示提升复杂环境下的单目深度估计精度

PromptMono: Cross Prompting Attention for Self-Supervised Monocular Depth Estimation in Challenging Environments

  • 引入可学习视觉提示捕捉不同场景特征
  • 在Oxford Robotcar和nuScenes上优于现有自监督方法
  • 适合需要跨环境泛化的自动驾驶深度感知任务

尽管理想条件下单目深度估计已取得进展,但在复杂环境中仍面临挑战。本文提出统一模型下的视觉提示学习方法,构建自监督框架PromptMono,通过一组可学习参数作为视觉提示以捕获特定领域知识。为融合提示信息至图像表征,设计新型门控交叉提示注意力(GCPA)模块,显著增强多样环境中的深度预测能力。在Oxford Robotcar与nuScenes数据集上的实验表明,所提方法性能优越。

原文摘要 · Abstract (English)

Considerable efforts have been made to improve monocular depth estimation under ideal conditions. However, in challenging environments, monocular depth estimation still faces difficulties. In this paper, we introduce visual prompt learning for predicting depth across different environments within a unified model, and present a self-supervised learning framework called PromptMono. It employs a set of learnable parameters as visual prompts to capture domain-specific knowledge. To integrate prompting information into image representations, a novel gated cross prompting attention (GCPA) module is proposed, which enhances the depth estimation in diverse conditions. We evaluate the proposed PromptMono on the Oxford Robotcar dataset and the nuScenes dataset. Experimental results demonstrate the superior performance of the proposed method.

单目深度视觉提示自监督自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。