用CLIP模型提升点云补全的空间精度,更准还原缺失部分。
Position-aware Guided Point Cloud Completion with CLIP Model
- 引入位置感知模块,通过加权映射增强缺失区域的空间信息。
- 构建图文点云三元组数据集,利用CLIP获取更丰富的细节表征。
- 在多个基准上超越现有方法,适合需要高精度3D重建的场景。
点云补全是为修复因设备缺陷或视角限制导致的几何与拓扑缺失而设计的任务。现有方法或仅依赖点云三维坐标进行补全,或结合已校准内参的图像来指导缺失部分的几何估计。尽管这些方法能直接预测完整点的位置并取得优异性能,但其提取的特征缺乏对缺失区域位置的细粒度描述。为此,本文提出一种快速高效的多模态扩展框架,通过引入位置感知模块,利用加权映射机制增强缺失区域的空间信息。同时,基于现有单模态点云补全数据集构建了点-文-图三元组数据集PCI-TI和MVP-TI,借助预训练的视觉语言模型CLIP为3D形状提供更丰富的细节信息,从而显著提升补全性能。大量定量与定性实验表明,该方法优于当前最先进点云补全技术。
原文摘要 · Abstract (English)
Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。