用光线为中心的模型生成高保真4D激光雷达数据,提升自动驾驶仿真真实度。
LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- 基于光线中心的时空注意力机制,建模点云沿射线的动态演化。
- 生成质量提升76%(MMD降低),重建准确率提高32%,预测误差减半。
- 支持可控生成,适合需要高精度场景模拟的研究与开发人员。
生成高保真且可控制的4D激光雷达数据对构建可扩展的自动驾驶仿真环境至关重要。该任务因传感器独特的球形几何、点云的时间稀疏性及动态场景复杂性而极具挑战。为此,我们提出LiSTAR,一种直接作用于传感器原生几何的新型生成世界模型。LiSTAR采用混合圆柱-球面(HCS)表示,通过减少笛卡尔网格中的量化伪影来保持数据保真度。为捕捉稀疏时间数据中的复杂动态,引入基于光线中心的时空注意力机制(START),显式建模各传感器射线上特征的演化过程,确保时间一致性。此外,针对可控生成,提出4D点云对齐的体素布局作为条件输入,并设计离散掩码生成型START(MaskSTART)框架,学习场景的紧凑分词表示,实现高效、高分辨率、布局引导的组合生成。大量实验验证了LiSTAR在4D激光雷达重建、预测和条件生成方面的顶尖性能:生成的MMD降低76%,重建交并比(IoU)提升32%,预测L1中位数误差下降50%。这一性能水平为构建逼真且可控的自动驾驶系统仿真提供了强大基础。
原文摘要 · Abstract (English)
Synthesizing high-fidelity and controllable 4D LiDAR data is crucial for creating scalable simulation environments for autonomous driving. This task is inherently challenging due to the sensor's unique spherical geometry, the temporal sparsity of point clouds, and the complexity of dynamic scenes. To address these challenges, we present LiSTAR, a novel generative world model that operates directly on the sensor's native geometry. LiSTAR introduces a Hybrid-Cylindrical-Spherical (HCS) representation to preserve data fidelity by mitigating quantization artifacts common in Cartesian grids. To capture complex dynamics from sparse temporal data, it utilizes a Spatio-Temporal Attention with Ray-Centric Transformer (START) that explicitly models feature evolution along individual sensor rays for robust temporal coherence. Furthermore, for controllable synthesis, we propose a novel 4D point cloud-aligned voxel layout for conditioning and a corresponding discrete Masked Generative START (MaskSTART) framework, which learns a compact, tokenized representation of the scene, enabling efficient, high-resolution, and layout-guided compositional generation. Comprehensive experiments validate LiSTAR's state-of-the-art performance across 4D LiDAR reconstruction, prediction, and conditional generation, with substantial quantitative gains: reducing generation MMD by a massive 76%, improving reconstruction IoU by 32%, and lowering prediction L1 Med by 50%. This level of performance provides a powerful new foundation for creating realistic and controllable autonomous systems simulations. Project link: https://ocean-luna.github.io/LiSTAR.gitub.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。