用语义体素引导扩散模型,生成大规模一致的驾驶场景。
SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation

- 基于语义体素网格的3D生成框架,存储彩色表面采样点。
- 通过局部体素扩散生成,支持大场景渐进式扩展与多视角一致渲染。
- 无需每场景优化,适合自动驾驶仿真与大规模场景生成。
大规模户外驾驶场景生成需要在多视角间保持一致的3D表示,并能扩展至大区域。现有方法或依赖从图像/视频模型蒸馏出的3D表示,损害几何一致性且渲染受限于训练视角;或仅适用于小规模3D场景或以物体为中心的生成。本文提出基于Σ-Voxfield网格的3D生成框架,该离散表示中每个占据体素存储固定数量的彩色表面样本。通过训练语义条件扩散模型,在局部体素邻域上操作并使用3D位置编码捕捉空间结构。利用重叠区域上的渐进式空间外推实现大场景扩展。最后,通过延迟渲染模块渲染生成的Σ-Voxfield网格,获得逼真图像,实现无需每场景优化的大规模多视角一致3D场景生成。大量实验表明,该方法可生成多样化的大规模城市户外场景,支持多种传感器配置和相机轨迹的逼真图像渲染,计算成本相比现有方法仍处于合理范围。
原文摘要 · Abstract (English)
Scalable generation of outdoor driving scenes requires 3D representations that remain consistent across multiple viewpoints and scale to large areas. Existing solutions either rely on image or video generative models distilled to 3D space, harming the geometric coherence and restricting the rendering to training views, or are limited to small-scale 3D scene or object-centric generation. In this work, we propose a 3D generative framework based on $Σ$-Voxfield grid, a discrete representation where each occupied voxel stores a fixed number of colorized surface samples. To generate this representation, we train a semantic-conditioned diffusion model that operates on local voxel neighborhoods and uses 3D positional encodings to capture spatial structure. We scale to large scenes via progressive spatial outpainting over overlapping regions. Finally, we render the generated $Σ$-Voxfield grid with a deferred rendering module to obtain photorealistic images, enabling large-scale multiview-consistent 3D scene generation without per-scene optimization. Extensive experiments show that our approach can generate diverse large-scale urban outdoor scenes, renderable into photorealistic images with various sensor configurations and camera trajectories while maintaining moderate computation cost compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。