arXiv:2510.02186cs.CVcs.LG2025-10中稿 · ICLR被引 1

用几何先验净化2D模型生成的3D点云,仅用1.5%数据达顶尖性能

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

  • 通过小网络从3D自监督教师中提炼几何先验,净化2D模型生成的3D特征
  • 在ScanNet、SUNRGB-D等基准上仅用1.5%训练数据即达领先效果
  • 适合缺乏大量标注3D数据的研究者,尤其适用于开放词汇3D分割任务

将2D视觉语言模型(VLM)特征迁移到3D语义分割面临持续权衡:直接投影导致预测噪声大、碎片化,强制几何一致性又需昂贵训练和大规模标注数据。我们指出,问题根源在于主流的分割-匹配范式无法调和2D语义与3D几何结构。实际上,几何线索并未在2D到3D转换中消失,而是隐藏在噪声和视图聚合的特征中。为此,我们提出GeoPurify,利用一个小型学生亲和力网络,结合3D自监督教师模型提取的几何先验,净化2D VLM生成的3D点特征。推理时,设计几何引导池化模块进一步去噪,确保语义与结构一致。得益于潜在几何信息和学习到的亲和力网络,GeoPurify有效缓解该权衡,在多个主流3D基准上实现或超越当前最优性能,仅使用约1.5%的训练数据。

原文摘要 · Abstract (English)

Recent attempts to transfer features from 2D Vision-Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields noisy and fragmented predictions, whereas enforcing geometric coherence necessitates costly training pipelines and large-scale annotated 3D data. We argue that this limitation stems from the dominant segmentation-and-matching paradigm, which fails to reconcile 2D semantics with 3D geometric structure. The geometric cues are not eliminated during the 2D-to-3D transfer but remain latent within the noisy and view-aggregated features. To exploit this property, we propose GeoPurify that applies a small Student Affinity Network to purify 2D VLM-generated 3D point features using geometric priors distilled from a 3D self-supervised teacher model. During inference, we devise a Geometry-Guided Pooling module to further denoise the point cloud and ensure the semantic and structural consistency. Benefiting from latent geometric information and the learned affinity network, GeoPurify effectively mitigates the trade-off and achieves superior data efficiency. Extensive experiments on major 3D benchmarks demonstrate that GeoPurify achieves or surpasses state-of-the-art performance while utilizing only about 1.5% of the training data.

3D分割几何先验数据高效开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。