用深度图引导的半监督方法提升RGB-D场景解析精度
DepthMatch: Semi-Supervised RGB-D Scene Parsing through Depth-Guided Regularization
- 通过补丁混叠增强挖掘图像纹理与空间特征关系
- 在NYUv2和KITTI上达顶尖性能,边界预测更精准
- 适合需要减少标注成本的场景解析研究者
RGB-D场景解析方法能有效捕捉环境的语义与几何特征,在极端天气和低光照条件下表现优异。然而现有方法多依赖需大量人工标注的监督训练,成本高昂。为此,我们提出DepthMatch,一种专为RGB-D场景解析设计的半监督学习框架。为充分利用无标签数据,提出互补补丁混叠增强,探索RGB-D图像对中纹理与空间特征的潜在关联;设计轻量级空间先验注入器,替代传统复杂融合模块,提升异质特征融合效率;引入深度引导边界损失,增强模型边界预测能力。实验表明,DepthMatch在室内外场景均具高适用性,在NYUv2数据集上达到领先水平,并在KITTI Semantics基准上排名第一。
原文摘要 · Abstract (English)
RGB-D scene parsing methods effectively capture both semantic and geometric features of the environment, demonstrating great potential under challenging conditions such as extreme weather and low lighting. However, existing RGB-D scene parsing methods predominantly rely on supervised training strategies, which require a large amount of manually annotated pixel-level labels that are both time-consuming and costly. To overcome these limitations, we introduce DepthMatch, a semi-supervised learning framework that is specifically designed for RGB-D scene parsing. To make full use of unlabeled data, we propose complementary patch mix-up augmentation to explore the latent relationships between texture and spatial features in RGB-D image pairs. We also design a lightweight spatial prior injector to replace traditional complex fusion modules, improving the efficiency of heterogeneous feature fusion. Furthermore, we introduce depth-guided boundary loss to enhance the model's boundary prediction capabilities. Experimental results demonstrate that DepthMatch exhibits high applicability in both indoor and outdoor scenes, achieving state-of-the-art results on the NYUv2 dataset and ranking first on the KITTI Semantics benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。