破解透明场景深度模糊,实现无需重训练的多层深度预测
Towards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
- 用拉普拉斯变换视觉提示提取预训练模型隐藏深度
- 零样本下实现多层深度估计,提升3D理解精度
- 适合需要鲁棒三维推理与视频深度建模的研究者
深度模糊是空间场景理解的根本挑战,尤其在透明场景中,单层深度估计无法捕捉完整三维结构。现有模型受限于确定性预测,忽略了真实世界中的多层深度。为此,我们提出从单预测到多假设空间基础模型的范式转变。首先构建了 exttt{MD-3k}基准,通过多层空间关系标签和新指标揭示专家与基础模型的深度偏差。为解决深度模糊,提出拉普拉斯视觉提示(LVP),一种无需训练的频域提示技术,通过拉普拉斯变换的RGB输入从预训练模型中提取隐含深度。将LVP推断深度与标准RGB估计融合,可在不重训练模型的前提下生成多层深度。大量实验验证了LVP在零样本多层深度估计中的有效性,显著提升几何条件驱动的视觉生成、3D引导的空间推理及时间一致的视频级深度推断性能。基准与代码将在https://github.com/Xiaohao-Xu/Ambiguity-in-Space公开。
原文摘要 · Abstract (English)
Depth ambiguity is a fundamental challenge in spatial scene understanding, especially in transparent scenes where single-depth estimates fail to capture full 3D structure. Existing models, limited to deterministic predictions, overlook real-world multi-layer depth. To address this, we introduce a paradigm shift from single-prediction to multi-hypothesis spatial foundation models. We first present \texttt{MD-3k}, a benchmark exposing depth biases in expert and foundational models through multi-layer spatial relationship labels and new metrics. To resolve depth ambiguity, we propose Laplacian Visual Prompting (LVP), a training-free spectral prompting technique that extracts hidden depth from pre-trained models via Laplacian-transformed RGB inputs. By integrating LVP-inferred depth with standard RGB-based estimates, our approach elicits multi-layer depth without model retraining. Extensive experiments validate the effectiveness of LVP in zero-shot multi-layer depth estimation, unlocking more robust and comprehensive geometry-conditioned visual generation, 3D-grounded spatial reasoning, and temporally consistent video-level depth inference. Our benchmark and code will be available at https://github.com/Xiaohao-Xu/Ambiguity-in-Space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。