arXiv:2510.11565cs.CV2025-10

一个模型搞定任意点云的点击和文字提示分割,跨场景通用。

SNAP: Towards Segmenting Anything in Any Point Cloud

  • 统一模型支持点选和文字提示,覆盖室内外及航拍数据
  • 在9个零样本测试中8个领先,文本提示任务全胜
  • 自动生成掩码匹配CLIP,适合需要高效标注的科研与工程

交互式3D点云分割可通过用户引导提示实现复杂3D场景的高效标注。然而,现有方法通常局限于单一领域(室内或室外)或单一交互形式(仅点选或仅文字提示)。此外,多数据集训练常导致负迁移,使模型缺乏泛化能力。为此,我们提出SNAP(Segment aNything in Any Point cloud),一个支持跨领域、多模态提示(点选与文字)的统一3D交互分割模型。通过在7个涵盖室内、室外和航拍环境的数据集上训练,并采用领域自适应归一化防止负迁移,实现跨域泛化。针对文字提示分割,自动生成掩码提案并匹配CLIP文本嵌入,支持全景与开放词汇分割。大量实验表明,SNAP在9个零样本空间提示基准中有8个达到当前最优性能,在所有5个文字提示基准上表现优异。结果证明,统一模型可媲美甚至超越专用领域模型,为大规模3D标注提供实用工具。

原文摘要 · Abstract (English)

Interactive 3D point cloud segmentation enables efficient annotation of complex 3D scenes through user-guided prompts. However, current approaches are typically restricted in scope to a single domain (indoor or outdoor), and to a single form of user interaction (either spatial clicks or textual prompts). Moreover, training on multiple datasets often leads to negative transfer, resulting in domain-specific tools that lack generalizability. To address these limitations, we present SNAP (Segment aNything in Any Point cloud), a unified model for interactive 3D segmentation that supports both point-based and text-based prompts across diverse domains. Our approach achieves cross-domain generalizability by training on 7 datasets spanning indoor, outdoor, and aerial environments, while employing domain-adaptive normalization to prevent negative transfer. For text-prompted segmentation, we automatically generate mask proposals without human intervention and match them against CLIP embeddings of textual queries, enabling both panoptic and open-vocabulary segmentation. Extensive experiments demonstrate that SNAP consistently delivers high-quality segmentation results. We achieve state-of-the-art performance on 8 out of 9 zero-shot benchmarks for spatial-prompted segmentation and demonstrate competitive results on all 5 text-prompted benchmarks. These results show that a unified model can match or exceed specialized domain-specific approaches, providing a practical tool for scalable 3D annotation. Project page is at, https://neu-vi.github.io/SNAP/

点云分割交互式分割多模态提示跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。