arXiv:2608.00678cs.CV2026-08

解决单目深度估计在相机翻滚下的失效问题,提升模型鲁棒性。

Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

论文配图:Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation
图 1 · 摘自论文原文
  • 引入不变深度约束,通过几何与空间推理任务训练模型
  • 在五个数据集上显著提升不同翻滚角度下的深度估计精度
  • 方法无需改动推理结构,适合部署到现有深度模型中

尽管单目深度估计取得进展,当前领先模型仍对相机翻滚敏感,轻微倾斜即导致深度预测严重退化。我们将其归因于被忽视的‘水平先验’——由于人类视觉偏好和拍摄习惯,训练图像多为近水平姿态,形成长尾分布偏差。虽有重加权数据增强和地平线校准等缓解手段,但效果有限。本文提出不变深度约束(ID-Constraint),一种训练时的监督策略,通过联合优化深度主干网络与一系列几何与空间推理任务,引导其学习旋转稳定的深度相关表征。辅助预测头在训练后移除,保持原推理结构不变。在五个基准数据集、四种翻滚设置下进行大量实验,验证了方法的有效性。

原文摘要 · Abstract (English)

Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can result in substantial degradation in depth estimations. We attribute this problem to a previously overlooked phenomenon, termed the Horizontal Prior, which is a manifestation of long-tailed distribution bias: most training images are captured in approximately horizontal orientations due to human visual preferences and photographic habits. While intuitive remedies such as re-balanced data augmentation and horizon leveling provide partial improvements, they fail to fully address the issue. In this paper, we introduce Invariant Depth Constraint (ID-Constraint), a training-time supervision strategy that improves roll robustness by fine-tuning and jointly regularizing the depth backbone with a series of geometric and spatial reasoning tasks. These auxiliary objectives encourage the backbone to learn rotation-stable, depth-relevant representations, while the auxiliary prediction heads are discarded after training, leaving the original inference architecture unchanged. Extensive experiments on five benchmark datasets across four roll settings demonstrate the effectiveness of the proposed method.

单目深度旋转鲁棒性几何约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。