arXiv:2606.29600cs.CVcs.AI2026-06中稿 · ECCV

揭示单目深度模型对多重几何结构的偏好差异,提出新评测基准。

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

论文配图:One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
图 1 · 摘自论文原文
  • 构建双层序数基准MD-3k,量化模型对深度层的选择偏好。
  • 主流模型在相同输入下呈现不同深度层预测,表明存在多种合理解释。
  • 无需训练的频谱提示可显著改变模型输出,最高达75.5%多层关系准确率。

真实的三维世界应包含分层几何结构,即一条相机射线可能穿过多个可见且几何合理的表面。然而,单目深度估计通常将每个像素简化为单一标量深度。透明场景使这种模糊性可测量:同一射线可能穿过前景玻璃并观测背景,导致监督目标成为标注、数据和训练的约定,而非场景固有的真实。学习到的预测器会暴露这一约定作为其深度层偏好。我们提出多层深度-3k(MD-3k),一个稀疏双层序数基准,用于衡量深度层偏好与多层空间关系准确性(ML-SRA)。在MD-3k上,领先的深度基础模型在标准RGB输入下表现出多样化的层偏好,说明相同分层几何可在不同模型中被不同解析。我们进一步发现,无需训练的拉普拉斯视觉提示(LVP)可显著改变某些冻结模型的预测层。最强的RGB/LVP组合DAv2-L达到75.5%的ML-SRA。这些结果表明,深度基础模型可能表达互补的几何假设,而标准RGB推理未能揭示。我们呼吁社区以模糊性感知视角重新思考深度监督与评估,将多种有效三维解释视为需测量、保留与表达的几何结构。

原文摘要 · Abstract (English)

A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces. Monocular depth estimation, however, reduces this structure to one scalar depth per pixel. Transparent scenes make this ambiguity measurable: the same ray can pass through foreground glass and observe the background, turning the supervised target into a convention of annotation, data, and training rather than a scene-intrinsic truth. A learned predictor exposes this convention as its depth-layer preference. We introduce MultiDepth-3k (MD-3k), a sparse two-layer ordinal benchmark for measuring depth-layer preference and multi-layer spatial relationship accuracy (ML-SRA). On MD-3k, leading depth foundation models exhibit diverse layer preferences under standard RGB input, showing that the same layered geometry can be resolved differently across models. We further find that Laplacian Visual Prompting (LVP), a training-free spectral input transformation, can substantially change the reported layer for certain frozen models. The strongest RGB/LVP pair, DAv2-L, reaches 75.5% ML-SRA. These results suggest that depth foundation models may express complementary geometric hypotheses that standard RGB inference leaves unexpressed. We invite the community to rethink depth supervision and evaluation through an ambiguity-aware lens, where multiple valid 3D interpretations are treated as geometric structure to be measured, preserved, and expressed.

深度估计几何模糊多层结构模型偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。