用3D布局桥接多种输入与激光点云,实现从文字到场景的可控生成。
LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs
- 用3D布局统一文本、图像、点云输入,生成语义与深度控制信号。
- 基于范围图的ControlNet实现像素级对齐,生成高质量点云。
- 支持从文字、图片、草图生成自驾车仿真环境,适合自动驾驶研发。
生成逼真且多样的激光雷达点云对自动驾驶仿真至关重要。尽管已有方法能根据用户输入生成点云,但因激光点云分布复杂而控制信号简单,难以兼顾高质量与多样化控制。为此,我们提出LiDARDraft,利用3D布局作为桥梁,连接多样条件输入与点云生成。3D布局可由文本、图像等任意输入快速生成。具体地,我们将文本、图像和点云统一表示为3D布局,并进一步转化为语义与深度控制信号。随后,采用基于范围图的ControlNet引导点云生成。该像素级对齐方法在可控点云生成上表现优异,支持‘从零开始仿真’,可从任意文本描述、图像或草图构建自驾车仿真环境。
原文摘要 · Abstract (English)
Generating realistic and diverse LiDAR point clouds is crucial for autonomous driving simulation. Although previous methods achieve LiDAR point cloud generation from user inputs, they struggle to attain high-quality results while enabling versatile controllability, due to the imbalance between the complex distribution of LiDAR point clouds and the simple control signals. To address the limitation, we propose LiDARDraft, which utilizes the 3D layout to build a bridge between versatile conditional signals and LiDAR point clouds. The 3D layout can be trivially generated from various user inputs such as textual descriptions and images. Specifically, we represent text, images, and point clouds as unified 3D layouts, which are further transformed into semantic and depth control signals. Then, we employ a rangemap-based ControlNet to guide LiDAR point cloud generation. This pixel-level alignment approach demonstrates excellent performance in controllable LiDAR point clouds generation, enabling "simulation from scratch", allowing self-driving environments to be created from arbitrary textual descriptions, images and sketches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。