arXiv:2507.06651cs.CV2025-07ICCV被引 6

用扩散模型先验提升图像到点云配准精度

Diff$^2$I2P: Differentiable Image-to-Point Cloud Registration with Diffusion Prior

  • 引入扩散模型作为跨模态先验,指导特征学习
  • 在7-Scenes上注册召回率提升超7%
  • 适合做高精度3D重建与机器人定位的研究者

学习跨模态对应关系对图像到点云(I2P)配准至关重要。现有方法多依赖度量学习强制模态间特征对齐,忽视了图像与点云间的固有模态差异,导致难以保证准确的跨模态对应。为此,受近期大型扩散模型跨模态生成成功的启发,我们提出Diff$^2$I2P,一个完全可微的I2P配准框架,利用新颖有效的扩散先验来弥合模态差距。具体地,提出控制侧得分蒸馏(CSD)技术,从深度条件扩散模型中蒸馏知识,直接优化预测变换。然而,由于对应关系检索和PnP求解器的不可微性,变换梯度无法反向传播至跨模态特征。为此,进一步提出可变形对应调优(DCT)模块,以可微方式估计对应关系,并使用可微PnP求解器进行变换估计。通过这两个设计,扩散模型作为强先验,引导图像与点云的跨模态特征学习,形成鲁棒对应关系,显著提升配准性能。大量实验表明,Diff$^2$I2P持续优于当前最优的I2P配准方法,在7-Scenes基准上注册召回率提升超过7%。

原文摘要 · Abstract (English)

Learning cross-modal correspondences is essential for image-to-point cloud (I2P) registration. Existing methods achieve this mostly by utilizing metric learning to enforce feature alignment across modalities, disregarding the inherent modality gap between image and point data. Consequently, this paradigm struggles to ensure accurate cross-modal correspondences. To this end, inspired by the cross-modal generation success of recent large diffusion models, we propose Diff$^2$I2P, a fully Differentiable I2P registration framework, leveraging a novel and effective Diffusion prior for bridging the modality gap. Specifically, we propose a Control-Side Score Distillation (CSD) technique to distill knowledge from a depth-conditioned diffusion model to directly optimize the predicted transformation. However, the gradients on the transformation fail to backpropagate onto the cross-modal features due to the non-differentiability of correspondence retrieval and PnP solver. To this end, we further propose a Deformable Correspondence Tuning (DCT) module to estimate the correspondences in a differentiable way, followed by the transformation estimation using a differentiable PnP solver. With these two designs, the Diffusion model serves as a strong prior to guide the cross-modal feature learning of image and point cloud for forming robust correspondences, which significantly improves the registration. Extensive experimental results demonstrate that Diff$^2$I2P consistently outperforms SoTA I2P registration methods, achieving over 7% improvement in registration recall on the 7-Scenes benchmark.

3D配准扩散模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。