用单图生成3D场景,靠扩散模型+深度预测实现高效重建。
Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
- 用预训练2D扩散模型和深度模型生成合成几何数据。
- 在KITTI-360和Waymo上达到或超越多视角监督方法性能。
- 适合动态场景重建,无需昂贵3D标注或多视角输入。
从单张图像进行体素化场景重建在自动驾驶和机器人等领域至关重要。现有方法虽表现优异,但通常依赖昂贵的3D真值或多视角监督。本文提出利用预训练的2D扩散模型与深度预测模型,从单张图像生成合成场景几何,并用于蒸馏一个前馈式场景重建模型。在具有挑战性的KITTI-360和Waymo数据集上的实验表明,该方法性能可媲美甚至超越依赖多视角监督的先进基线,尤其在动态场景下展现出独特优势。
原文摘要 · Abstract (English)
Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achieve impressive results, but generally require expensive 3D ground truth or multi-view supervision. We propose to leverage pre-trained 2D diffusion models and depth prediction models to generate synthetic scene geometry from a single image. This can then be used to distill a feed-forward scene reconstruction model. Our experiments on the challenging KITTI-360 and Waymo datasets demonstrate that our method matches or outperforms state-of-the-art baselines that use multi-view supervision, and offers unique advantages, for example regarding dynamic scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。