用手机双像素数据实现高精度深度估计,突破传统方法局限
DiFuse-Net: RGB and Dual-Pixel Depth Estimation using Window Bi-directional Parallax Attention and Cross-modal Transfer Learning
- 设计窗口双向视差注意力机制,精准捕捉手机摄像头微小视差
- 在真实数据集上达到优于传统双目和双像素基线的方法性能
- 提出跨模态迁移学习,利用海量RGB-D数据弥补双像素数据不足
深度估计对智能系统至关重要,广泛应用于自动驾驶与增强现实。传统双目和主动传感器存在成本高、功耗大、鲁棒性差的问题,而现代摄像头普遍采用的双像素(DP)技术提供了更优替代方案。本文提出DiFuse-Net,一种解耦式多模态网络,专门用于分离处理RGB图像与双像素数据以进行深度估计。其核心为窗口双向视差注意力机制(WBiPAM),能有效捕获智能手机小光圈下细微的双像素视差特征。同时采用独立编码器提取RGB上下文信息,并进行特征融合提升预测精度。为解决高质量双像素深度数据稀缺问题,提出跨模态迁移学习(CmTL)机制,利用现有大规模RGB-D数据集进行预训练。实验表明该方法显著优于基于双像素和双目的基线模型。此外,本文构建了新的高质量真实世界RGB-DP-D训练数据集——双相机双像素(DCDP)数据集,通过创新的对称双摄硬件、立体校准与矫正流程及AI立体视差估计方法实现。
原文摘要 · Abstract (English)
Depth estimation is crucial for intelligent systems, enabling applications from autonomous navigation to augmented reality. While traditional stereo and active depth sensors have limitations in cost, power, and robustness, dual-pixel (DP) technology, ubiquitous in modern cameras, offers a compelling alternative. This paper introduces DiFuse-Net, a novel modality decoupled network design for disentangled RGB and DP based depth estimation. DiFuse-Net features a window bi-directional parallax attention mechanism (WBiPAM) specifically designed to capture the subtle DP disparity cues unique to smartphone cameras with small aperture. A separate encoder extracts contextual information from the RGB image, and these features are fused to enhance depth prediction. We also propose a Cross-modal Transfer Learning (CmTL) mechanism to utilize large-scale RGB-D datasets in the literature to cope with the limitations of obtaining large-scale RGB-DP-D dataset. Our evaluation and comparison of the proposed method demonstrates its superiority over the DP and stereo-based baseline methods. Additionally, we contribute a new, high-quality, real-world RGB-DP-D training dataset, named Dual-Camera Dual-Pixel (DCDP) dataset, created using our novel symmetric stereo camera hardware setup, stereo calibration and rectification protocol, and AI stereo disparity estimation method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。