arXiv:2601.03660cs.CV2026-01被引 2

提出多模态点云补全框架,提升真实场景下的泛化能力

MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding

  • 融合点云、图像和文本,设计模态丢弃策略增强鲁棒性
  • 在百万级数据集上实现优于基线的补全精度和真实场景泛化
  • 适合需要跨模态、高鲁棒性的3D重建研究者使用

点云补全旨在从受限视角和遮挡导致的不完整观测中恢复完整的三维几何结构。现有基于学习的方法,包括基于3D卷积神经网络、点云和Transformer的方法,在合成基准上表现强劲。然而,由于模态限制、可扩展性和生成能力不足,其在新物体和真实场景中的泛化能力仍面临挑战。本文提出MGPC,一种统一架构的通用多模态点云补全框架,整合点云、RGB图像和文本信息。MGPC引入创新的模态丢弃策略、基于Transformer的融合模块和新型渐进式生成器,以提升鲁棒性、可扩展性和几何建模能力。我们进一步构建了自动数据生成流程,建立包含超过1,000个类别和一百万个训练样本的大型基准MGPC-1M。在MGPC-1M和真实世界数据上的大量实验表明,所提方法持续优于已有基线,并在真实条件下展现出强泛化能力。

原文摘要 · Abstract (English)

Point cloud completion aims to recover complete 3D geometry from partial observations caused by limited viewpoints and occlusions. Existing learning-based works, including 3D Convolutional Neural Network (CNN)-based, point-based, and Transformer-based methods, have achieved strong performance on synthetic benchmarks. However, due to the limitations of modality, scalability, and generative capacity, their generalization to novel objects and real-world scenarios remains challenging. In this paper, we propose MGPC, a generalizable multimodal point cloud completion framework that integrates point clouds, RGB images, and text within a unified architecture. MGPC introduces an innovative modality dropout strategy, a Transformer-based fusion module, and a novel progressive generator to improve robustness, scalability, and geometric modeling capability. We further develop an automatic data generation pipeline and construct MGPC-1M, a large-scale benchmark with over 1,000 categories and one million training pairs. Extensive experiments on MGPC-1M and in-the-wild data demonstrate that the proposed method consistently outperforms prior baselines and exhibits strong generalization under real-world conditions.

点云补全多模态泛化能力3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。