arXiv:2503.19947cs.CVcs.AI2025-03中稿 · IntelliSys 2026

让预训练视觉模型学会通用深度理解,无需微调即可提升多任务性能。

Vanishing Depth: Training Generalized Depth Adapters with Sinusoidal Depth Preprocessing for Pretrained RGB Encoders

  • 用正弦深度编码与自监督训练扩展预训练模型,融合度量深度信息。
  • 在SUN-RGBD上达到56.05 mIoU,优于现有深度感知与多模态模型。
  • 支持无深度输入场景,可借助单像素或单目估计激活深度特征提取。

通用度量深度理解对精准视觉引导机器人至关重要,但当前最先进的视觉编码器无法支持。为此,我们提出一种自监督训练方法,通过添加深度适配器,将度量深度信息融入预训练RGB编码器的联合隐空间中,不干扰原有的RGB特征提取。结合正弦深度编码,该深度适配器实现了泛化性强、对深度密度与分布不变的特征提取。在分割、姿态估计和深度补全等一系列相关下游任务中,我们的深度适配器显著提升了多种通用RGB基线性能,且无需微调。最重要的是,在SUN-RGBD分割任务中取得56.05 mIoU,超越了现有的深度感知与多模态编码器。当无深度输入时,可通过空映射、单像素深度线索或单目深度估计激活深度感知特征提取,应用于后续任务。

原文摘要 · Abstract (English)

Generalized metric depth understanding is critical for precise vision-guided robotics, which current state-of-the-art (SOTA) vision-encoders do not support. To address this, we propose a self-supervised training approach that extends pretrained RGB encoders with a depth adapter to incorporate and align metric depth into a combined latent space without interfering with the pretrained RGB feature extraction. In combination with our sinusoidal depth encoding, the depth adapter enables generalized and robust depth density and distribution invariant feature extraction. Our depth adapters improve a wide set of generalized RGB baselines across a spectrum of relevant RGBD downstream tasks in segmentation, pose estimation, and depth completion -- without the necessity of finetuning. Most importantly, we achieve 56.05 mIoU in the SUN-RGBD segmentation, while outperforming SOTA depth-aware and multi-modal encoders in our experiments. When no depth is present, one can activate our depth adapter with an empty map, use single pixel depth clues, or monocular depth estimation to include the depth aware feature extraction into subsequent downstream tasks.

深度感知视觉编码器自监督学习机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。