arXiv:2511.06033cs.CVcs.AI2025-11被引 2

通过时空谱联合学习提升深度补全精度,解决传感器缺失数据问题。

S2ML: Spatio-Spectral Mutual Learning for Depth Completion

  • 融合空间与频域特征,利用幅度和相位谱差异设计专用融合模块。
  • 在NYU-Depth V2和SUN RGB-D上分别领先CFormer 0.828dB和0.834dB。
  • 适合做高精度深度补全的科研与工业应用,尤其关注物理特性建模。

基于时差(TOF)或结构光的RGB-D相机获取的原始深度图像常因弱反射、边界阴影和伪影导致深度值不完整,限制了下游视觉任务的应用。现有方法多在图像域进行深度补全,却忽略了原始深度图像的物理特性。研究发现,无效深度区域会改变其频率分布模式。为此,本文提出时空谱互学习框架(S2ML),协同利用空间域与频域的优势实现深度补全。具体地,针对幅度谱与相位谱的差异设计专用谱融合模块,并在统一嵌入空间中建模空间与频域特征的局部与全局相关性。通过渐进式互表示与精炼机制,网络能充分挖掘互补的物理先验以提升补全精度。大量实验表明,S2ML在NYU-Depth V2和SUN RGB-D数据集上分别优于当前最优方法CFormer 0.828dB和0.834dB。

原文摘要 · Abstract (English)

The raw depth images captured by RGB-D cameras using Time-of-Flight (TOF) or structured light often suffer from incomplete depth values due to weak reflections, boundary shadows, and artifacts, which limit their applications in downstream vision tasks. Existing methods address this problem through depth completion in the image domain, but they overlook the physical characteristics of raw depth images. It has been observed that the presence of invalid depth areas alters the frequency distribution pattern. In this work, we propose a Spatio-Spectral Mutual Learning framework (S2ML) to harmonize the advantages of both spatial and frequency domains for depth completion. Specifically, we consider the distinct properties of amplitude and phase spectra and devise a dedicated spectral fusion module. Meanwhile, the local and global correlations between spatial-domain and frequency-domain features are calculated in a unified embedding space. The gradual mutual representation and refinement encourage the network to fully explore complementary physical characteristics and priors for more accurate depth completion. Extensive experiments demonstrate the effectiveness of our proposed S2ML method, outperforming the state-of-the-art method CFormer by 0.828 dB and 0.834 dB on the NYU-Depth V2 and SUN RGB-D datasets, respectively.

深度补全频域分析多域融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。