arXiv:2508.16579cs.CV2025-08被引 2

融合单目深度先验,提升iToF深度精度与视野

Towards High-Precision Depth Sensing via Monocular-Aided iToF and RGB Integration

  • 将窄视场iToF深度图重投影至宽视场RGB坐标系,实现像素级对齐
  • 双编码器融合网络结合结构约束,提升深度边缘清晰度与分辨率
  • 适用于复杂场景下的高精度深度感知,适合机器人视觉等应用

本文提出一种新颖的iToF-RGB融合框架,以解决间接飞行时间(iToF)深度感知固有的低空间分辨率、有限视场(FoV)及复杂场景中的结构失真问题。首先通过精确的几何标定与对齐模块,将窄视场iToF深度图重投影到宽视场RGB坐标系,确保模态间像素级对应。随后采用双编码器融合网络,联合提取重投影后的iToF深度与RGB图像的互补特征,并利用单目深度先验恢复细粒度结构细节,实现深度超分辨率。通过整合跨模态结构线索与深度一致性约束,该方法显著提升深度精度、边缘锐度并实现无缝视场扩展。在合成与真实世界数据集上的大量实验表明,所提框架在准确性、结构一致性和视觉质量方面均显著优于现有先进方法。

原文摘要 · Abstract (English)

This paper presents a novel iToF-RGB fusion framework designed to address the inherent limitations of indirect Time-of-Flight (iToF) depth sensing, such as low spatial resolution, limited field-of-view (FoV), and structural distortion in complex scenes. The proposed method first reprojects the narrow-FoV iToF depth map onto the wide-FoV RGB coordinate system through a precise geometric calibration and alignment module, ensuring pixel-level correspondence between modalities. A dual-encoder fusion network is then employed to jointly extract complementary features from the reprojected iToF depth and RGB image, guided by monocular depth priors to recover fine-grained structural details and perform depth super-resolution. By integrating cross-modal structural cues and depth consistency constraints, our approach achieves enhanced depth accuracy, improved edge sharpness, and seamless FoV expansion. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed framework significantly outperforms state-of-the-art methods in terms of accuracy, structural consistency, and visual quality.

深度估计多模态融合iToF深度超分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。