arXiv:2503.18393cs.CV2025-03AAAI被引 5

用伪深度图替代真实深度,提升复杂室内场景语义分割效果

PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes

  • 通过多张伪深度图聚合生成统一深度模态
  • 在NYUv2上达6.98 mIoU提升,超越现有方法
  • 基于扩散模型提取特征,适合多模态分割任务

RGB与深度模态的融合显著提升了复杂室内场景的语义分割精度,其中来自RGB-D相机的深度数据起关键作用。然而,采集RGB-D数据集比纯RGB更昂贵,且存在传感器定位、数据缺失和噪声等对齐难题。相比之下,高精度深度估计算法生成的伪深度(PD)可摆脱对专用深度传感器和对齐流程的依赖,同时提供有效深度信息,在语义分割中展现巨大潜力。为探索伪深度用于分割的可行性,本文设计了基于RGB-PD的分割流水线,提出伪深度聚合模块(PDAM),将多张伪深度图融合为单一模态,便于集成至其他RGB-D分割方法。此外,预训练扩散模型作为强大特征提取器适用于RGB分割任务,但多模态扩散分割方法尚未被充分研究。因此,本文提出伪深度扩散模型(PDDM),采用大规模文本-图像扩散模型作为特征提取器,并设计简单有效的融合策略整合伪深度。在NYUv2和SUNRGB-D数据集上的大量实验表明,伪深度能有效提升分割性能,所提PDDM达到当前最优水平,在NYUv2上优于其他方法6.98 mIoU,SUNRGB-D上提升2.11 mIoU。

原文摘要 · Abstract (English)

The integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more expensive than an RGB dataset due to the need for specialized depth sensors. Aligning depth and RGB images also poses challenges due to sensor positioning and issues like missing data and noise. In contrast, Pseudo Depth (PD) from high-precision depth estimation algorithms can eliminate the dependence on RGB-D sensors and alignment processes, as well as provide effective depth information and show significant potential in semantic segmentation. Therefore, to explore the practicality of utilizing pseudo depth instead of real depth for semantic segmentation, we design an RGB-PD segmentation pipeline to integrate RGB and pseudo depth and propose a Pseudo Depth Aggregation Module (PDAM) for fully exploiting the informative clues provided by the diverse pseudo depth maps. The PDAM aggregates multiple pseudo depth maps into a single modality, making it easily adaptable to other RGB-D segmentation methods. In addition, the pre-trained diffusion model serves as a strong feature extractor for RGB segmentation tasks, but multi-modal diffusion-based segmentation methods remain unexplored. Therefore, we present a Pseudo Depth Diffusion Model (PDDM) that adopts a large-scale text-image diffusion model as a feature extractor and a simple yet effective fusion strategy to integrate pseudo depth. To verify the applicability of pseudo depth and our PDDM, we perform extensive experiments on the NYUv2 and SUNRGB-D datasets. The experimental results demonstrate that pseudo depth can effectively enhance segmentation performance, and our PDDM achieves state-of-the-art performance, outperforming other methods by +6.98 mIoU on NYUv2 and +2.11 mIoU on SUNRGB-D.

语义分割伪深度扩散模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。