用自然语言生成可编辑的4D激光雷达序列,提升自动驾驶仿真精度。
Learning to Generate 4D LiDAR Sequences
- 通过语言指令生成激光雷达场景图,分三支扩散模型合成物体布局、轨迹与形状。
- 在nuScenes数据集上实现最优的感知质量、可控性与时序一致性。
- 支持物体插入/移动等编辑操作,适合自动驾驶数据增强与仿真研究。
尽管生成式世界模型已在视频和占用数据合成方面取得进展,但激光雷达生成仍不充分,而其对精准3D感知至关重要。将生成扩展至4D激光雷达数据面临可控性、时序稳定性和评估难题。我们提出LiDARCrafter,一个统一框架,可将自由文本指令转化为可编辑的激光雷达序列。指令被解析为本体中心场景图,由三分支扩散模型生成物体布局、轨迹与形状;范围图像扩散模型生成初始扫描,自回归模块将其扩展为时序连贯序列。显式布局设计支持物体级编辑(如插入或重定位)。为实现公平评估,我们提供EvalSuite基准,涵盖场景、物体和序列级指标。在nuScenes数据集上,LiDARCrafter在保真度、可控性和时序一致性上达到当前最佳表现,为基于激光雷达的仿真与数据增强奠定基础。
原文摘要 · Abstract (English)
While generative world models have advanced video and occupancy-based data synthesis, LiDAR generation remains underexplored despite its importance for accurate 3D perception. Extending generation to 4D LiDAR data introduces challenges in controllability, temporal stability, and evaluation. We present LiDARCrafter, a unified framework that converts free-form language into editable LiDAR sequences. Instructions are parsed into ego-centric scene graphs, which a tri-branch diffusion model transforms into object layouts, trajectories, and shapes. A range-image diffusion model generates the initial scan, and an autoregressive module extends it into a temporally coherent sequence. The explicit layout design further supports object-level editing, such as insertion or relocation. To enable fair assessment, we provide EvalSuite, a benchmark spanning scene-, object-, and sequence-level metrics. On nuScenes, LiDARCrafter achieves state-of-the-art fidelity, controllability, and temporal consistency, offering a foundation for LiDAR-based simulation and data augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。