arXiv:2607.15600cs.CV2026-07

用立体图像对齐信息,让单目深度模型学会准确的尺度判断。

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

论文配图:Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth
图 1 · 摘自论文原文
  • 通过校准立体图像令牌,将多视角几何先验迁移到单目模型
  • 在ETH3D和DIODE上零样本深度估计精度显著提升
  • 不依赖推理时的多视角输入,兼容主流ViT模型

单目深度基础模型在多样环境中展现出强大的泛化能力,但仍在度量深度估计上存在困难。这一限制源于单视角推理固有的尺度模糊性,导致即使相对几何准确,尺度预测仍不匹配。相比之下,近期多视角基础模型利用跨视角线索学习鲁棒的场景级几何与一致尺度,但在单图推理时这些优势通常消失,因缺乏显式几何约束导致性能下降。为此,我们提出一种新框架,将多视角模型的尺度感知几何先验迁移到单目深度基础模型中。具体地,引入基于校准立体图像的视差蒸馏(EpiDistill),利用校准立体图像令牌使单视角预测模型保持视差注意力模式并维持几何一致性,无需推理时提供多视角输入。实验表明,该方法显著提升了零样本度量深度估计性能,尤其在尺度对齐至关重要的ETH3D和DIODE等挑战性数据集上表现突出。此外,该方法具有模型无关性,能持续提升当前最优的ViT-based模型(如UniDepthV2和DepthPro)性能。

原文摘要 · Abstract (English)

Monocular depth foundation models have demonstrated remarkable generalization capabilities across diverse environments. However, they continue to struggle with metric depth estimation in diverse environments. This limitation stems from the inherent scale ambiguity of single-view inference, leading to misaligned scale predictions even when the relative geometry is accurate. Conversely, recent multi-view foundation models leverage cross-view cues to learn robust scene-level geometry and consistent scale. Yet, these benefits typically vanish during single-image inference, as the absence of explicit geometric constraints causes performance to degrade. To bridge this gap, we propose a novel framework that transfers the scale-aware geometric priors of multi-view models into monocular depth foundation models. Specifically, we introduce an Epipolar Distillation (EpiDistill), an approach utilizing Rectified Stereo Tokens, which enables the single-view prediction model to retain epipolar attention patterns and maintain geometric consistency without requiring multi-view inputs at inference. Experimental results demonstrate that our method significantly improves zero-shot metric depth estimation, particularly on challenging datasets like ETH3D and DIODE where scale alignment is critical. Furthermore, our approach is model-agnostic, consistently boosting the performance of state-of-the-art ViT-based models, including UniDepthV2 and DepthPro.

单目深度几何蒸馏立体匹配尺度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。