用视觉模型的语义先验提升激光雷达注册鲁棒性,无需摄像头
CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

- 从视觉大模型蒸馏语义先验到点云,增强对密度和视角变化的适应力
- 在KITTI等数据集上达97.7%以上成功率,稀疏扫描下仍保持97.3%精度
- 无需相机输入或后处理优化,支持跨传感器零样本泛化
基于学习的全局点云注册虽取得显著进展,但依赖几何表示,对点云密度、扫描模式、视角和传感器特性变化敏感。本文提出CVSD-Reg框架,将视觉基础模型中的视觉语义先验蒸馏至激光雷达表征中。第一阶段,采用冻结的DINOv2作为教师模型,通过对比蒸馏与球面流形对齐,使点云学生模型保留教师嵌入空间的超球几何结构;自监督InfoNCE一致性与软SE(3)不变性进一步提升视角鲁棒性。第二阶段,通过对应学习、密度感知点丢弃增强及端到端位姿优化,实现注册适配。仅需单个检查点,即可在单传感器和零样本跨传感器场景下通用,且推理阶段完全无需相机。在KITTI、nuScenes和HeLiPR上,严格成功率达97.7%、99.0%和99.3%,包括稀疏16束Velodyne扫描下的97.3%;相比现有几何注册方法最高提升44.0个百分点,无需相机输入或后续ICP精修。
原文摘要 · Abstract (English)
Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations. In Stage 1, a Point Transformer V3 student learns from a frozen DINOv2 teacher through contrastive distillation and spherical-manifold alignment, which preserves the hyperspherical geometry of the teacher embedding space. Self-supervised InfoNCE consistency and soft $\mathrm{SE}(3)$ invariance further encourage viewpoint-robust descriptors. In Stage 2, the distilled representation is adapted to registration through correspondence learning, density-aware point-dropout augmentation, and end-to-end pose optimization. With a single checkpoint, CVSD-Reg generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference. On KITTI, nuScenes, and HeLiPR, CVSD-Reg achieves strict success rate ([email protected]\,m/$1^\circ$) of 97.7$\%$, 99.0$\%$, and 99.3$\%$, respectively, including 97.3$\%$ on sparse 16-beam Velodyne scans. It outperforms state-of-the-art geometric registration methods by up to 44.0 percentage points without requiring camera inputs or post-hoc ICP refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。