无需标注数据,让图像与点云自动对齐,提升自动驾驶感知精度。
Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration
- 双路径对比学习构建图像与点云的联合语义几何嵌入空间。
- 粗到精注册流程在KITTI上提升23.7%、nuScenes上提升37.9%。
- 动态训练机制平衡多任务损失,适合自动驾驶多模态系统。
连接2D与3D传感器模态对自动驾驶系统的鲁棒感知至关重要。然而,由于纹理丰富但深度模糊的图像与稀疏但度量精确的点云之间存在语义-几何鸿沟,且现有方法易陷入局部最优,图像到点云(I2P)配准仍具挑战。为此,我们提出CrossI2P,一种统一跨模态学习与两阶段配准的自监督框架。首先,通过双路径对比学习构建几何-语义融合嵌入空间,实现无标注的双向对齐。其次,采用粗到精注册策略:全局阶段通过联合同模态上下文与跨模态交互建模建立超点-超像素对应,随后进行几何约束的点级精修以实现高精度配准。第三,引入梯度归一化的动态训练机制,平衡特征对齐、对应关系优化与位姿估计的损失。大量实验表明,CrossI2P在KITTI Odometry基准上优于现有方法23.7%,在nuScenes上提升37.9%,显著提升准确率与鲁棒性。
原文摘要 · Abstract (English)
Bridging 2D and 3D sensor modalities is critical for robust perception in autonomous systems. However, image-to-point cloud (I2P) registration remains challenging due to the semantic-geometric gap between texture-rich but depth-ambiguous images and sparse yet metrically precise point clouds, as well as the tendency of existing methods to converge to local optima. To overcome these limitations, we introduce CrossI2P, a self-supervised framework that unifies cross-modal learning and two-stage registration in a single end-to-end pipeline. First, we learn a geometric-semantic fused embedding space via dual-path contrastive learning, enabling annotation-free, bidirectional alignment of 2D textures and 3D structures. Second, we adopt a coarse-to-fine registration paradigm: a global stage establishes superpoint-superpixel correspondences through joint intra-modal context and cross-modal interaction modeling, followed by a geometry-constrained point-level refinement for precise registration. Third, we employ a dynamic training mechanism with gradient normalization to balance losses for feature alignment, correspondence refinement, and pose estimation. Extensive experiments demonstrate that CrossI2P outperforms state-of-the-art methods by 23.7% on the KITTI Odometry benchmark and by 37.9% on nuScenes, significantly improving both accuracy and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。