arXiv:2603.12721cs.CVcs.AI2026-03被引 12

融合2D图像与3D点云,提升复杂场景下的点云配准精度

CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration

  • 结合2D图像上下文与3D几何特征,构建混合注意力网络
  • 在3DMatch和3DLoMatch上显著提升配准准确率与鲁棒性
  • 适用于低重叠、噪声干扰等真实复杂场景,适合视觉-几何任务研究者

点云配准是3D计算机视觉与几何深度学习的基础任务,对大规模3D重建、增强现实和场景理解至关重要。然而,现有基于学习的方法在存在数据不完整、传感器噪声和低重叠区域的复杂真实场景中性能下降。为此,我们提出CMHANet——一种跨模态混合注意力网络。该方法融合2D图像的丰富上下文信息与3D点云的几何细节,生成全面且稳健的特征表示。同时,引入基于对比学习的新优化函数,强化几何一致性,显著提升模型对噪声和部分观测的鲁棒性。我们在3DMatch和具有挑战性的3DLoMatch数据集上评估了CMHANet,零样本测试在TUM RGB-D SLAM数据集上验证了其对未见领域的泛化能力。实验结果表明,该方法在配准精度和整体鲁棒性方面均显著优于现有技术。代码已开源。

原文摘要 · Abstract (English)

Robust point cloud registration is a fundamental task in 3D computer vision and geometric deep learning, essential for applications such as large-scale 3D reconstruction, augmented reality, and scene understanding. However, the performance of established learning-based methods often degrades in complex, real world scenarios characterized by incomplete data, sensor noise, and low overlap regions. To address these limitations, we propose CMHANet, a novel Cross-Modal Hybrid Attention Network. Our method integrates the fusion of rich contextual information from 2D images with the geometric detail of 3D point clouds, yielding a comprehensive and resilient feature representation. Furthermore, we introduce an innovative optimization function based on contrastive learning, which enforces geometric consistency and significantly improves the model's robustness to noise and partial observations. We evaluated CMHANet on the 3DMatch and the challenging 3DLoMatch datasets. \rev{Additionally, zero-shot evaluations on the TUM RGB-D SLAM dataset verify the model's generalization capability to unseen domains.} The experimental results demonstrate that our method achieves substantial improvements in both registration accuracy and overall robustness, outperforming current techniques. We also release our code in \href{https://github.com/DongXu-Zhang/CMHANet}{https://github.com/DongXu-Zhang/CMHANet}.

点云配准跨模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。