用扩散模型生成带精确几何的3D驾驶场景,支持可控、真实且结构一致的视角合成。
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- 结合2D图像先验与评分蒸馏,生成代理几何与环境表示。
- 支持地图布局条件下的提示引导生成,实现高保真纹理与结构。
- 适合需要真实几何一致性与轨迹多样性的自动驾驶训练场景。
大规模场景数据对机器人学习中的训练与测试至关重要。神经重建方法虽能从传感器数据中重建大尺度物理接地的户外场景,但受限于静态环境与有限的场景控制能力,其场景与轨迹多样性受制于原始采集数据。相比之下,基于图像或视频扩散模型的生成方法虽具备较高控制性,却缺乏几何接地与因果一致性。本文旨在弥合这一差距,提出一种直接生成具精确几何结构的大规模3D驾驶场景的方法,支持因果性新视角合成、物体持续存在及显式3D几何估计。该方法结合代理几何与环境表示的生成,以及来自已学2D图像先验的评分蒸馏。实验表明,该方法可实现高可控性,支持以地图布局为条件的提示引导生成,产出具有高保真纹理与结构、几何一致的复杂驾驶场景。
原文摘要 · Abstract (English)
Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these methods have baked-in static environments and only allow for limited scene control -- they are functionally constrained in scene and trajectory diversity by the captures from which they are reconstructed. In contrast, generating driving data with recent image or video diffusion models offers control, however, at the cost of geometry grounding and causality. In this work, we aim to bridge this gap and present a method that directly generates large-scale 3D driving scenes with accurate geometry, allowing for causal novel view synthesis with object permanence and explicit 3D geometry estimation. The proposed method combines the generation of a proxy geometry and environment representation with score distillation from learned 2D image priors. We find that this approach allows for high controllability, enabling the prompt-guided geometry and high-fidelity texture and structure that can be conditioned on map layouts -- producing realistic and geometrically consistent 3D generations of complex driving scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。