构建最大规模驾驶场景占用数据集,实现多模态高保真生成。
Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

- 基于时空解耦架构,统一生成语义占用、视频与点云。
- 在100万+帧数据上实现4D动态占用的高精度时序建模。
- 适合自动驾驶感知与规划评估,支持多传感器真实仿真。
驾驶场景生成是自动驾驶的关键领域,支持感知与规划评估等下游应用。近年来,以占用为中心的方法通过跨帧与跨模态的一致性条件实现了最先进性能,但其表现严重依赖标注的占用数据,而此类数据仍极为稀缺。为解决这一问题,我们构建了迄今最大的语义占用数据集Nuplan-Occ,基于广泛使用的Nuplan基准构建。其规模与多样性不仅支持大规模生成建模,也赋能自动驾驶下游任务。基于该数据集,我们提出统一框架,联合生成高质量语义占用、多视角视频与激光雷达点云。方法采用时空解耦架构,支持4D动态占用的高保真空间扩展与时间预测。为弥合模态差距,提出两项新策略:基于高斯点阵的稀疏点图渲染方法,提升多视角视频生成质量;以及传感器感知嵌入策略,显式建模激光雷达特性以实现真实多激光雷达模拟。大量实验表明,本方法在生成保真度与可扩展性上优于现有方法,并验证其在下游任务中的实用价值。
原文摘要 · Abstract (English)
Driving scene generation is a critical domain for autonomous driving, enabling downstream applications, including perception and planning evaluation. Occupancy-centric methods have recently achieved state-of-the-art results by offering consistent conditioning across frames and modalities; however, their performance heavily depends on annotated occupancy data, which still remains scarce. To overcome this limitation, we curate Nuplan-Occ, the largest semantic occupancy dataset to date, constructed from the widely used Nuplan benchmark. Its scale and diversity facilitate not only large-scale generative modeling but also autonomous driving downstream applications. Based on this dataset, we develop a unified framework that jointly synthesizes high-quality semantic occupancy, multi-view videos, and LiDAR point clouds. Our approach incorporates a spatio-temporal disentangled architecture to support high-fidelity spatial expansion and temporal forecasting of 4D dynamic occupancy. To bridge modal gaps, we further propose two novel techniques: a Gaussian splatting-based sparse point map rendering strategy that enhances multi-view video generation, and a sensor-aware embedding strategy that explicitly models LiDAR sensor properties for realistic multi-LiDAR simulation. Extensive experiments demonstrate that our method achieves superior generation fidelity and scalability compared to existing approaches, and validates its practical value in downstream tasks. Repo: https://github.com/Arlo0o/UniScene-Unified-Occupancy-centric-Driving-Scene-Generation/tree/v2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。