arXiv:2503.00793cs.CVcs.AI2025-03ICRA被引 5

用几何引导对比学习,让多光谱图像统一深度估计更可靠高效。

Bridging Spectral-wise and Multi-spectral Depth Estimation via Geometry-guided Contrastive Learning

  • 通过几何引导对比学习对齐多光谱特征空间,实现跨波段共享表征。
  • 设计可插拔融合模块,选择性聚合多光谱特征,提升预测鲁棒性。
  • 单网络支持多光谱融合与光谱不变估计,兼顾效率与灵活性。

将深度估计网络部署于真实世界需具备应对多种恶劣条件的强鲁棒性,以保障自主系统安全可靠。为此,许多自动驾驶车辆采用多模态传感器系统,包括RGB相机、NIR相机、热成像相机、LiDAR或雷达。现有方法主要分为两类:模态独立推理(灵活但内存低效、不可靠)和多模态融合推理(可靠性高但需专用架构)。本文提出一种名为“对齐-融合”的有效方案,用于从多光谱图像中进行深度估计。在对齐阶段,利用几何线索最小化全局与局部特征的对比损失,对齐不同光谱波段的嵌入空间,学习跨波段共享表征。在融合阶段,训练一个可插拔的特征融合模块,有选择地聚合多光谱特征,实现可靠且鲁棒的预测结果。基于该方法,单一深度网络可同时实现光谱不变与多光谱融合的深度估计,兼顾可靠性、内存效率与灵活性。

原文摘要 · Abstract (English)

Deploying depth estimation networks in the real world requires high-level robustness against various adverse conditions to ensure safe and reliable autonomy. For this purpose, many autonomous vehicles employ multi-modal sensor systems, including an RGB camera, NIR camera, thermal camera, LiDAR, or Radar. They mainly adopt two strategies to use multiple sensors: modality-wise and multi-modal fused inference. The former method is flexible but memory-inefficient, unreliable, and vulnerable. Multi-modal fusion can provide high-level reliability, yet it needs a specialized architecture. In this paper, we propose an effective solution, named align-and-fuse strategy, for the depth estimation from multi-spectral images. In the align stage, we align embedding spaces between multiple spectrum bands to learn shareable representation across multi-spectral images by minimizing contrastive loss of global and spatially aligned local features with geometry cue. After that, in the fuse stage, we train an attachable feature fusion module that can selectively aggregate the multi-spectral features for reliable and robust prediction results. Based on the proposed method, a single-depth network can achieve both spectral-invariant and multi-spectral fused depth estimation while preserving reliability, memory efficiency, and flexibility.

深度估计多光谱对比学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。