通过三路融合与颜色感知注意力,提升图像到点云配准精度
TFCT-I2P: Three stream fusion network with color aware transformer for image-to-point cloud registration
- 设计三路融合网络,联合图像颜色与点云结构特征
- 引入颜色感知变换器,有效缓解局部错配问题
- 在多个数据集上超越当前最优方法,适合三维视觉任务
随着人工智能技术的发展,图像到点云配准(I2P)取得了显著进展。然而,点云(三维)与图像(二维)特征在维度上的差异仍带来巨大挑战,主要表现为难以利用一种模态的特征增强另一种,导致潜在空间中的特征对齐困难。为此,我们提出一种名为TFCT-I2P的图像到点云配准方法。首先,设计三路融合网络(TFN),将图像的颜色信息与点云的结构信息融合,促进双模态特征对齐;其次,为有效缓解引入颜色信息后产生的局部块级错配,提出颜色感知变压器(CAT);最后,在7Scenes、RGB-D Scenes V2、ScanNet V2及自收集数据集上进行大量实验。结果表明,相较于现有最先进方法,TFCT-I2P在内点率上提升1.5%,特征匹配召回率提升0.4%,配准召回率提升5.4%。因此,我们认为所提方法推动了I2P配准技术的发展。
原文摘要 · Abstract (English)
Along with the advancements in artificial intelligence technologies, image-to-point-cloud registration (I2P) techniques have made significant strides. Nevertheless, the dimensional differences in the features of points cloud (three-dimension) and image (two-dimension) continue to pose considerable challenges to their development. The primary challenge resides in the inability to leverage the features of one modality to augment those of another, thereby complicating the alignment of features within the latent space. To address this challenge, we propose an image-to-point-cloud method named as TFCT-I2P. Initially, we introduce a Three-Stream Fusion Network (TFN), which integrates color information from images with structural information from point clouds, facilitating the alignment of features from both modalities. Subsequently, to effectively mitigate patch-level misalignments introduced by the inclusion of color information, we design a Color-Aware Transformer (CAT). Finally, we conduct extensive experiments on 7Scenes, RGB-D Scenes V2, ScanNet V2, and a self-collected dataset. The results demonstrate that TFCT-I2P surpasses state-of-the-art methods by 1.5% in Inlier Ratio, 0.4% in Feature Matching Recall, and 5.4% in Registration Recall. Therefore, we believe that the proposed TFCT-I2P contributes to the advancement of I2P registration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。