arXiv:2507.22825cs.CV2025-07ICCV被引 22

用深度图引导扩散模型,分步重建单视角场景中的每个物体。

DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion

  • 分步生成物体再组合,避免整体重建的复杂性
  • 在合成与真实数据上均达顶尖性能,泛化能力强
  • 深度信息贯穿训练与推理,充分挖掘几何细节

我们提出 DepR,一种基于深度图的单视角场景重建框架,将实例级扩散模型融入组合式建模范式。不同于以往方法仅在推理时用深度图估计物体布局、未能充分利用其丰富几何信息的情况,DepR 在训练和推理阶段均引入深度引导条件,有效将形状先验编码至扩散模型。推理时,深度图进一步指导 DDIM 采样与布局优化,提升重建结果与输入图像的一致性。尽管仅在有限合成数据上训练,DepR 在合成与真实世界数据集上的评估中均达到当前最优表现,展现出强大的泛化能力。

原文摘要 · Abstract (English)

We propose DepR, a depth-guided single-view scene reconstruction framework that integrates instance-level diffusion within a compositional paradigm. Instead of reconstructing the entire scene holistically, DepR generates individual objects and subsequently composes them into a coherent 3D layout. Unlike previous methods that use depth solely for object layout estimation during inference and therefore fail to fully exploit its rich geometric information, DepR leverages depth throughout both training and inference. Specifically, we introduce depth-guided conditioning to effectively encode shape priors into diffusion models. During inference, depth further guides DDIM sampling and layout optimization, enhancing alignment between the reconstruction and the input image. Despite being trained on limited synthetic data, DepR achieves state-of-the-art performance and demonstrates strong generalization in single-view scene reconstruction, as shown through evaluations on both synthetic and real-world datasets.

3D重建扩散模型深度引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。