arXiv:2607.01900cs.CV2026-07

用全局结构先验提升单目深度弱视差区域的精度

FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation

论文配图:FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation
图 1 · 摘自论文原文
  • 融合双像素深度与单目基础模型的全局结构先验
  • 在低对比度区域提升结构保真度,误差降低18.3%
  • 适合需要高精度深度图的自动驾驶与机器人应用

双像素(DP)成像通过子孔径视差实现单相机测距,但极小的有效基线导致视差可观测性差,在无纹理、低对比或下采样区域易出现结构退化和深度失效。现有方法主要依赖局部视差线索,在视差信号弱或模糊时不可靠。为此,我们提出FoundDP,一个统一框架,将可度量的双像素深度与来自单目深度基础模型的全局结构先验相结合。该方法通过双像素深度保持度量尺度,并利用视觉变换器(ViT)特征恢复弱视差区域的结构一致性。为确保在双像素成像条件下可靠度量引导,我们识别并缓解了由双像素散焦模糊引起的ViT表示退化,通过特征对齐实现稳定度量引导的深度估计。在合成与真实世界双像素基准上的大量实验表明,FoundDP在结构保真度和度量准确性上均表现更优,尤其在视差可观测性降低时仍具显著增益。代码将公开于:https://github.com/EchoLighting/FoundDP

原文摘要 · Abstract (English)

Dual-pixel (DP) imaging enables metric depth estimation from a single camera using sub-aperture disparity. However, the extremely small effective baseline limits disparity observability, leading to structural degradation and depth failure in textureless, low-contrast, or downsampled regions. Existing DP-based methods rely primarily on local disparity cues and therefore become unreliable when disparity signals are weak or ambiguous. To address this limitation, we propose \emph{FoundDP}, a unified framework that integrates metric DP depth with global structural priors from a monocular depth foundation model. Our method preserves metric scale through DP-derived depth and leverages Vision Transformer (ViT) features to restore structural consistency in weak-disparity regions. To ensure reliable metric guidance under DP imaging conditions, we identify and mitigate ViT representation degradation induced by DP defocus blur via ViT feature alignment, enabling stable metric-guided depth estimation. Extensive experiments on synthetic and real-world DP benchmarks show that FoundDP delivers superior performance, with consistent gains in structural fidelity and metric accuracy, especially under reduced disparity observability. Code will be available at: https://github.com/EchoLighting/FoundDP

深度估计双像素结构先验视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。