用自监督模型做奖励,让3D点云自动分物体,无需人工标注。
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation

- 用超点迭代合并+强化学习,自动发现复杂场景中的物体。
- 在多个数据集上超越基线,零样本和长尾场景表现优异。
- 适合做无标注3D分割,尤其适用于大规模场景建模。
我们解决复杂场景点云中无场景级人工标注的3D物体分割难题。现有方法通常只能识别简单物体,主要因学习过程中缺乏足够的物体先验。本文提出FoundObj框架,包含基于超点的物体发现代理,通过语义与几何奖励模块引导,逐步合并相邻超点。这两个模块融合自监督2D/3D基础模型提供的语义与几何先验,为物体发现代理提供互补反馈,借助强化学习实现多类别物体的鲁棒识别。在多个基准上的大量实验表明,本方法持续优于现有基线。尤其在零样本与长尾场景下展现强泛化能力,凸显其在可扩展、无标签3D物体分割中的潜力。
原文摘要 · Abstract (English)
We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during training. Existing methods are typically constrained to identifying simple objects, primarily due to insufficient object priors in the learning process. In this paper, we present FoundObj, a novel framework featuring a superpoint-based object discovery agent that incrementally merges suitable neighboring superpoints, guided by our innovative semantic and geometric reward modules. These modules synergistically leverage semantic and geometric priors from self-supervised 2D/3D foundation models, providing complementary feedback to the object discovery agent and enabling robust identification of multi-class objects through reinforcement learning. Extensive experiments on diverse benchmarks demonstrate that our approach consistently outperforms existing baselines. Notably, our method exhibits strong generalization in zero-shot and long-tail scenarios, underscoring its potential for scalable, label-free 3D object segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。