通过一致性2D掩码追踪,实现无需人工标注的3D实例分割
Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
- 基于时序一致的2D掩码跟踪,解决多帧独立处理导致的粒度不一致问题
- 在ScanNet和SUNRGB-D上达到最新水平,开放词汇能力显著提升
- 适合缺乏标注数据的3D场景理解任务,尤其适用于视频序列数据
3D实例分割对现实应用至关重要。为避免高昂的人工标注成本,现有方法尝试将基础模型生成的2D掩码迁移到3D空间以生成伪标签。然而,由于视频帧独立处理,常导致分割粒度不一致和3D伪标签冲突,降低最终分割精度。为此,本文提出一种粒度一致的自动2D掩码追踪方法,保持帧间时序对应关系,消除冲突伪标签。结合三阶段课程学习框架,从单视角碎片化数据逐步过渡到统一多视角标注,最终实现全局一致的完整场景监督。该结构化学习流程使模型能逐步适应一致性更高的伪标签,从而从初始碎片化且矛盾的2D先验中鲁棒地提炼出一致的3D表征。实验表明,该方法可有效生成一致且准确的3D分割结果,在标准基准(ScanNet、SUNRGB-D)上达到当前最优性能,并具备良好的开放词汇能力。
原文摘要 · Abstract (English)
3D instance segmentation is an important task for real-world applications. To avoid costly manual annotations, existing methods have explored generating pseudo labels by transferring 2D masks from foundation models to 3D. However, this approach is often suboptimal since the video frames are processed independently. This causes inconsistent segmentation granularity and conflicting 3D pseudo labels, which degrades the accuracy of final segmentation. To address this, we introduce a Granularity-Consistent automatic 2D Mask Tracking approach that maintains temporal correspondences across frames, eliminating conflicting pseudo labels. Combined with a three-stage curriculum learning framework, our approach progressively trains from fragmented single-view data to unified multi-view annotations, ultimately globally coherent full-scene supervision. This structured learning pipeline enables the model to progressively expose to pseudo-labels of increasing consistency. Thus, we can robustly distill a consistent 3D representation from initially fragmented and contradictory 2D priors. Experimental results demonstrated that our method effectively generated consistent and accurate 3D segmentations. Furthermore, the proposed method achieved state-of-the-art results on standard benchmarks and open-vocabulary ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。