arXiv:2503.01257cs.CV2025-03CVPR被引 5

用频率自适应融合提升手机dToF深度图的清晰度与连贯性

SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective Fusion

  • 通过多帧融合与自适应卷积核选择,解决稀疏深度图的模糊问题
  • 在TartanAir和Dynamic Replica上实现最优精度与低闪烁表现
  • 适合移动端3D感知应用,尤其对实时深度补全有显著提升

轻量级直接飞行时间(dToF)传感器适用于移动设备上的3D感知。然而,由于紧凑设备的制造限制及成像物理原理,dToF深度图存在稀疏且噪声大的问题。本文提出一种名为SVDC的新颖视频深度补全方法,通过融合稀疏dToF数据与对应RGB图像实现增强。采用多帧融合策略以缓解由稀疏成像带来的空间模糊。为解决多帧融合中帧间错位导致的边缘与背景混合问题,提出自适应频率选择融合(AFSF)模块,自动选择卷积核大小进行特征融合。AFSF利用通道-空间增强注意力(CSEA)模块增强特征,并生成注意力图作为融合权重,有效恢复边缘细节并抑制平滑区域高频噪声。为进一步提升时序一致性,设计跨窗口一致性损失,确保不同窗口间预测一致,显著减少闪烁现象。SVDC在TartanAir与Dynamic Replica数据集上达到最佳准确率与一致性。代码已开源:https://github.com/Lan1eve/SVDC。

原文摘要 · Abstract (English)

Lightweight direct Time-of-Flight (dToF) sensors are ideal for 3D sensing on mobile devices. However, due to the manufacturing constraints of compact devices and the inherent physical principles of imaging, dToF depth maps are sparse and noisy. In this paper, we propose a novel video depth completion method, called SVDC, by fusing the sparse dToF data with the corresponding RGB guidance. Our method employs a multi-frame fusion scheme to mitigate the spatial ambiguity resulting from the sparse dToF imaging. Misalignment between consecutive frames during multi-frame fusion could cause blending between object edges and the background, which results in a loss of detail. To address this, we introduce an adaptive frequency selective fusion (AFSF) module, which automatically selects convolution kernel sizes to fuse multi-frame features. Our AFSF utilizes a channel-spatial enhancement attention (CSEA) module to enhance features and generates an attention map as fusion weights. The AFSF ensures edge detail recovery while suppressing high-frequency noise in smooth regions. To further enhance temporal consistency, We propose a cross-window consistency loss to ensure consistent predictions across different windows, effectively reducing flickering. Our proposed SVDC achieves optimal accuracy and consistency on the TartanAir and Dynamic Replica datasets. Code is available at https://github.com/Lan1eve/SVDC.

深度补全dToF视频生成多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。