arXiv:2412.02336cs.CV2024-12ICCV被引 6

让模型看穿遮挡,预测物体被遮住部分的深度。

Amodal Depth Anything: Amodal Depth Estimation in the Wild

  • 用可见部分推断被遮挡区域的相对深度,提升真实场景泛化能力。
  • 在新数据集ADIW上实现69.5%的精度提升,超越当前最佳方法。
  • 结合确定性与生成式模型,可生成多样且合理的遮挡深度结构。

非可见深度估计旨在预测场景中被遮挡(不可见)物体部分的深度。该任务探讨模型能否仅凭可见线索有效感知被遮挡区域的几何信息。以往方法主要依赖合成数据集,聚焦于度量深度估计,受限于域偏移和可扩展性问题,难以推广至真实世界。本文提出一种面向真实场景的非可见深度估计新范式,专注于相对深度预测以增强模型跨自然图像的泛化能力。我们构建了一个大规模新数据集ADIW,采用可扩展的流水线,利用分割数据集和合成技术生成深度图,并通过尺度-偏移对齐策略优化与融合深度预测,确保标注一致性。为解决该任务,提出两种互补框架:基于Depth Anything V2的确定性模型Amodal-DAV2,以及融合条件流匹配原理的生成式模型Amodal-DepthFM。所提框架仅需对大型预训练模型进行最小修改,即可实现高质量的非可见深度预测。实验验证设计合理性,表明模型能生成多样化、合理的遮挡深度结构。在ADIW数据集上,方法相较此前最优成果实现69.5%的准确率提升。

原文摘要 · Abstract (English)

Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive the geometry of occluded regions based on visible cues. Prior methods primarily rely on synthetic datasets and focus on metric depth estimation, limiting their generalization to real-world settings due to domain shifts and scalability challenges. In this paper, we propose a novel formulation of amodal depth estimation in the wild, focusing on relative depth prediction to improve model generalization across diverse natural images. We introduce a new large-scale dataset, Amodal Depth In the Wild (ADIW), created using a scalable pipeline that leverages segmentation datasets and compositing techniques. Depth maps are generated using large pre-trained depth models, and a scale-and-shift alignment strategy is employed to refine and blend depth predictions, ensuring consistency in ground-truth annotations. To tackle the amodal depth task, we present two complementary frameworks: Amodal-DAV2, a deterministic model based on Depth Anything V2, and Amodal-DepthFM, a generative model that integrates conditional flow matching principles. Our proposed frameworks effectively leverage the capabilities of large pre-trained models with minimal modifications to achieve high-quality amodal depth predictions. Experiments validate our design choices, demonstrating the flexibility of our models in generating diverse, plausible depth structures for occluded regions. Our method achieves a 69.5% improvement in accuracy over the previous SoTA on the ADIW dataset.

深度估计遮挡推理生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。