让激光雷达场景生成更可控,可精确布局物体位置。
La La LiDAR: Large-Scale Layout Generation from LiDAR Data

- 用场景图扩散+关系感知条件,分步生成结构化场景
- 在Waymo和nuScenes数据集上生成效果领先,下游感知任务表现优
- 适合自动驾驶仿真与安全验证,支持自定义物体布局
可控生成真实激光雷达场景对自动驾驶和机器人应用至关重要。现有基于扩散模型的方法虽能生成高保真点云,但缺乏对前景物体和空间关系的显式控制,限制了其在场景仿真与安全验证中的应用。为此,我们提出大型布局引导的激光雷达生成模型(La La LiDAR),采用语义增强的场景图扩散与关系感知的上下文条件机制,实现结构化布局生成,并引入前景感知的控制注入完成完整场景生成。该方法可实现物体位置的定制化控制,同时保证空间与语义一致性。为支持结构化生成,我们构建了两个大规模激光雷达场景图数据集:Waymo-SG 和 nuScenes-SG,以及新的布局合成评估指标。大量实验表明,La La LiDAR 在激光雷达生成与下游感知任务中均达到领先性能,确立了可控三维场景生成的新基准。
原文摘要 · Abstract (English)
Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit control over foreground objects and spatial relationships, limiting their usefulness for scenario simulation and safety validation. To address these limitations, we propose Large-scale Layout-guided LiDAR generation model ("La La LiDAR"), a novel layout-guided generative framework that introduces semantic-enhanced scene graph diffusion with relation-aware contextual conditioning for structured LiDAR layout generation, followed by foreground-aware control injection for complete scene generation. This enables customizable control over object placement while ensuring spatial and semantic consistency. To support our structured LiDAR generation, we introduce Waymo-SG and nuScenes-SG, two large-scale LiDAR scene graph datasets, along with new evaluation metrics for layout synthesis. Extensive experiments demonstrate that La La LiDAR achieves state-of-the-art performance in both LiDAR generation and downstream perception tasks, establishing a new benchmark for controllable 3D scene generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。