不依赖图像的点云补全方法表现优于当前最先进模型。
A Strong View-Free Baseline Approach for Single-View Image Guided Point Cloud Completion
- 仅用部分点云输入,通过多分支注意力网络实现补全。
- 在ShapeNet-ViPC数据集上超越现有最先进方法。
- 揭示了图像引导对点云补全的必要性存疑,适合多模态学习研究者。
单视图图像引导点云补全(SVIPC)任务旨在利用单张图像和部分点云重建完整点云。尽管已有研究证明该多模态方法有效,但图像引导的根本必要性仍缺乏深入探讨。为此,我们提出一种强基线方法,基于注意力驱动的多分支编码器-解码器网络,仅使用部分点云作为输入,实现无视角依赖(view-free)。其分层自融合机制通过交叉注意力与自注意力层,有效整合多路特征,增强几何结构建模能力。在ShapeNet-ViPC数据集上的大量实验与消融研究显示,该无图像视角框架性能优于当前最先进的SVIPC方法。我们的发现为多模态学习在SVIPC中的发展提供了新视角。演示代码将公开于https://github.com/Zhang-VISLab。
原文摘要 · Abstract (English)
The single-view image guided point cloud completion (SVIPC) task aims to reconstruct a complete point cloud from a partial input with the help of a single-view image. While previous works have demonstrated the effectiveness of this multimodal approach, the fundamental necessity of image guidance remains largely unexamined. To explore this, we propose a strong baseline approach for SVIPC based on an attention-based multi-branch encoder-decoder network that only takes partial point clouds as input, view-free. Our hierarchical self-fusion mechanism, driven by cross-attention and self-attention layers, effectively integrates information across multiple streams, enriching feature representations and strengthening the networks ability to capture geometric structures. Extensive experiments and ablation studies on the ShapeNet-ViPC dataset demonstrate that our view-free framework performs superiorly to state-of-the-art SVIPC methods. We hope our findings provide new insights into the development of multimodal learning in SVIPC. Our demo code will be available at https://github.com/Zhang-VISLab.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。