arXiv:2508.14036cs.CVcs.AI2025-08被引 12

用2D交互提示实现3D零件分割,无需文本或完整标注。

GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation

  • 将3D分割转为多视角2D掩码预测,通过点击/框引导零件选择。
  • 在PartObjaverse-Tiny和PartNetE上达到顶尖性能,优于优化与前馈方法。
  • 支持细粒度控制,适合需要精确交互的3D理解场景。

我们提出GeoSAM2,一种可提示控制的3D零件分割框架,将任务转化为多视图2D掩码预测。对于无纹理物体,从预设视角渲染法向量图和点云图,并接受简单的2D提示(点击或框)以指导零件选择。这些提示由共享SAM2主干网络处理,该网络通过LoRA和残差几何融合增强,实现在保留预训练先验的同时进行视图特定推理。预测的掩码被反投影至物体并跨视图聚合。本方法无需文本提示、每形状优化或全3D标签,即可实现细粒度、零件级控制。相比全局聚类或基于尺度的方法,提示具有显式、空间定位且可解释性高。在PartObjaverse-Tiny和PartNetE上取得当前最优的类无关性能,超越依赖优化的慢速流程和快速但粗糙的前馈方法。结果揭示新范式:将3D分割与SAM2对齐,利用交互式2D输入,在对象级零件理解中实现可控性与精度。

原文摘要 · Abstract (English)

We introduce GeoSAM2, a prompt-controllable framework for 3D part segmentation that casts the task as multi-view 2D mask prediction. Given a textureless object, we render normal and point maps from predefined viewpoints and accept simple 2D prompts - clicks or boxes - to guide part selection. These prompts are processed by a shared SAM2 backbone augmented with LoRA and residual geometry fusion, enabling view-specific reasoning while preserving pretrained priors. The predicted masks are back-projected to the object and aggregated across views. Our method enables fine-grained, part-specific control without requiring text prompts, per-shape optimization, or full 3D labels. In contrast to global clustering or scale-based methods, prompts are explicit, spatially grounded, and interpretable. We achieve state-of-the-art class-agnostic performance on PartObjaverse-Tiny and PartNetE, outperforming both slow optimization-based pipelines and fast but coarse feedforward approaches. Our results highlight a new paradigm: aligning the paradigm of 3D segmentation with SAM2, leveraging interactive 2D inputs to unlock controllability and precision in object-level part understanding.

3D分割交互式多视图几何融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。