用单帧+动作指令生成超长距3D交通场景,突破自动驾驶仿真规模瓶颈。
OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models
- 基于占用世界模型与布局生成器,仅凭初始帧和动作序列生成长时序场景。
- 可稳定生成超3000帧连续画面,构建4公里以上3D占用图,比前代提升80倍。
- 生成数据可直接用于预训练模型,零样本性能领先传统仿真工具11%~22.1%。
数据驱动的自动驾驶仿真长期依赖预录驾驶日志或高精地图等先验,严重限制了可扩展性,使开放式生成局限于已有数据集规模。为此,我们提出首个基于占用世界模型的3D仿真器OccSim。OccSim无需持续日志或高精地图,仅需单一初始帧和未来自车动作序列,即可稳定生成超过3000帧连续画面,实现长达4公里以上的3D占用图连续构建,相比此前最先进模型在稳定生成长度上提升超80倍。OccSim由两个模块构成:基于W-DiT的静态占用世界模型与布局生成器。前者通过显式引入已知刚体变换设计,实现超长时序静态环境生成;后者基于合成道路拓扑,动态填充具有响应能力的前景代理。该设计支持大规模、多样化的仿真流生成。大量实验表明其下游价值:直接从OccSim采集的数据可用于预训练4D语义占用预测模型,在未见数据上达到最高67%的零样本性能,优于以往基于资产的仿真器11%;当数据集规模扩大至5倍时,零样本性能提升至约74%,对资产类仿真器的优势扩大至22.1%。
原文摘要 · Abstract (English)
Data-driven autonomous driving simulation has long been constrained by its heavy reliance on pre-recorded driving logs or spatial priors, such as HD maps. This fundamental dependency severely limits scalability, restricting open-ended generation capabilities to the finite scale of existing collected datasets. To break this bottleneck, we present OccSim, the first occupancy world model-driven 3D simulator. OccSim obviates the requirement for continuous logs or HD maps; conditioned only on a single initial frame and a sequence of future ego-actions, it can stably generate over 3,000 continuous frames, enabling the continuous construction of large-scale 3D occupancy maps spanning over 4 kilometers for simulation. This represents an >80x improvement in stable generation length over previous state-of-the-art occupancy world models. OccSim is powered by two modules: W-DiT based static occupancy world model and the Layout Generator. W-DiT handles the ultra-long-horizon generation of static environments by explicitly introducing known rigid transformations in architecture design, while the Layout Generator populates the dynamic foreground with reactive agents based on the synthesized road topology. With these designs, OccSim can synthesize massive, diverse simulation streams. Extensive experiments demonstrate its downstream utility: data collected directly from OccSim can pre-train 4D semantic occupancy forecasting models to achieve up to 67% zero-shot performance on unseen data, outperforming previous asset-based simulator by 11%. When scaling the OccSim dataset to 5x the size, the zero-shot performance increases to about 74%, while the improvement over asset-based simulators expands to 22.1%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。