用自然语言生成动态4D激光雷达场景,支持精细编辑与流畅时序。
LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
- 通过语言解析生成以自我为中心的场景图,控制物体结构、运动和几何。
- 在nuScenes数据集上实现最高保真度、可控性与时序一致性。
- 开源代码与基准测试,适合自动驾驶仿真与数据增强研究者。
生成式世界模型已成为自动驾驶的关键数据引擎,但现有工作多聚焦于视频或占据网格,忽视了激光雷达的独特属性。将激光雷达生成扩展至动态4D世界建模面临可控性、时序一致性和评估标准化的挑战。为此,我们提出LiDARCrafter,一个统一的4D激光雷达生成与编辑框架。给定自由形式的自然语言输入,系统解析指令为以自我为中心的场景图,作为三分支扩散网络的条件,分别生成物体结构、运动轨迹与几何。这些结构化条件支持多样且细粒度的场景编辑。此外,自回归模块生成具有平滑过渡的时序一致4D激光雷达序列。为支持标准化评估,我们建立了一个涵盖场景、物体和序列层面的综合性基准。在nuScenes数据集上的实验表明,LiDARCrafter在所有层级均达到最先进的保真度、可控性与时序一致性,为数据增强与仿真开辟新路径。代码与基准已开源。
原文摘要 · Abstract (English)
Generative world models have become essential data engines for autonomous driving, yet most existing efforts focus on videos or occupancy grids, overlooking the unique LiDAR properties. Extending LiDAR generation to dynamic 4D world modeling presents challenges in controllability, temporal coherence, and evaluation standardization. To this end, we present LiDARCrafter, a unified framework for 4D LiDAR generation and editing. Given free-form natural language inputs, we parse instructions into ego-centric scene graphs, which condition a tri-branch diffusion network to generate object structures, motion trajectories, and geometry. These structured conditions enable diverse and fine-grained scene editing. Additionally, an autoregressive module generates temporally coherent 4D LiDAR sequences with smooth transitions. To support standardized evaluation, we establish a comprehensive benchmark with diverse metrics spanning scene-, object-, and sequence-level aspects. Experiments on the nuScenes dataset using this benchmark demonstrate that LiDARCrafter achieves state-of-the-art performance in fidelity, controllability, and temporal consistency across all levels, paving the way for data augmentation and simulation. The code and benchmark are released to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。