用3D重建增强2D图像语义分割,提升弱监督效果
Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation
- 通过2D视频重建3D点云,提供几何辅助监督信号
- 在稀疏标注下实现2-7%的性能提升,优于现有方法
- 适合资源有限但需高精度分割的应用场景
我们提出Rewis3d框架,利用前馈式3D重建技术显著提升2D图像上弱监督语义分割的性能。获取密集像素级标注仍是训练分割模型的主要瓶颈。稀疏标注可作为高效弱监督替代方案,但仍存在性能差距。为此,我们引入一种新方法,将3D场景重建作为辅助监督信号。核心思想是:从2D视频中恢复的3D几何结构能为稀疏标注提供强传播线索。具体地,采用双学生-教师架构,在2D图像与重建3D点云间强制语义一致性,并使用先进的前馈重建生成可靠几何监督。大量实验表明,Rewis3d在稀疏监督下达到当前最优性能,相比现有方法提升2-7%,且无需额外标签或推理开销。
原文摘要 · Abstract (English)
We present Rewis3d, a framework that leverages recent advances in feed-forward 3D reconstruction to significantly improve weakly supervised semantic segmentation on 2D images. Obtaining dense, pixel-level annotations remains a costly bottleneck for training segmentation models. Alleviating this issue, sparse annotations offer an efficient weakly-supervised alternative. However, they still incur a performance gap. To address this, we introduce a novel approach that leverages 3D scene reconstruction as an auxiliary supervisory signal. Our key insight is that 3D geometric structure recovered from 2D videos provides strong cues that can propagate sparse annotations across entire scenes. Specifically, a dual student-teacher architecture enforces semantic consistency between 2D images and reconstructed 3D point clouds, using state-of-the-art feed-forward reconstruction to generate reliable geometric supervision. Extensive experiments demonstrate that Rewis3d achieves state-of-the-art performance in sparse supervision, outperforming existing approaches by 2-7% without requiring additional labels or inference overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。