arXiv:2501.09947cs.CV2025-01中稿 · TIP被引 16

用神经表面表示实现无标注物体分割,多视角图像自动对齐更准。

Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation

  • 通过双神经表面模块建模复杂场景几何结构
  • 在多个基准上优于基于NeRF的自监督方法,接近有监督水平
  • 适合无标注数据场景下的高精度物体分割研究

自监督物体分割(SOS)旨在无需标注的情况下分割物体。在多相机输入条件下,可利用各视角间的结构、纹理和几何一致性实现细粒度分割。为此,我们提出基于神经表面表示的自监督物体分割(Surface-SOS),通过多视角图像构建3D表面表示,实现每个视图的物体分割。为建模高质量复杂场景表面,设计了一种新颖的场景表示方案,将场景分解为两个互补的神经表示模块,均采用带符号距离函数(SDF)。此外,Surface-SOS通过引入粗分割掩码作为额外输入,利用多视角未标注图像优化单视图分割结果。据我们所知,Surface-SOS是首个利用神经表面表示打破对大量标注数据和强约束依赖的自监督方法。这些约束通常要求目标物体位于静态背景或依赖视频中的时间监督。在标准基准数据集LLFF、CO3D、BlendedMVS、TUM及多个真实场景上的大量实验表明,Surface-SOS始终生成比其基于NeRF的对应方法更精细的物体掩码,并显著超越有监督单视图基线。代码已公开于:https://github.com/zhengxyun/Surface-SOS。

原文摘要 · Abstract (English)

Self-supervised Object Segmentation (SOS) aims to segment objects without any annotations. Under conditions of multi-camera inputs, the structural, textural and geometrical consistency among each view can be leveraged to achieve fine-grained object segmentation. To make better use of the above information, we propose Surface representation based Self-supervised Object Segmentation (Surface-SOS), a new framework to segment objects for each view by 3D surface representation from multi-view images of a scene. To model high-quality geometry surfaces for complex scenes, we design a novel scene representation scheme, which decomposes the scene into two complementary neural representation modules respectively with a Signed Distance Function (SDF). Moreover, Surface-SOS is able to refine single-view segmentation with multi-view unlabeled images, by introducing coarse segmentation masks as additional input. To the best of our knowledge, Surface-SOS is the first self-supervised approach that leverages neural surface representation to break the dependence on large amounts of annotated data and strong constraints. These constraints typically involve observing target objects against a static background or relying on temporal supervision in videos. Extensive experiments on standard benchmarks including LLFF, CO3D, BlendedMVS, TUM and several real-world scenes show that Surface-SOS always yields finer object masks than its NeRF-based counterparts and surpasses supervised single-view baselines remarkably. Code is available at: https://github.com/zhengxyun/Surface-SOS.

自监督分割神经表面多视角几何无标注学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。