提出可联合生成图像与点云的端到端框架,解决自动驾驶多模态生成难题。
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
- 通过相机与鸟瞰图间双向转换模块,实现2D与3D模态联合生成。
- 引入深度预测分支,提升从图像空间还原为鸟瞰空间的准确性。
- 支持时序预测,适用于真实自动驾驶场景的未来世界建模。
生成模型在自动驾驶的相机图像或激光雷达点云生成与预测方面已取得显著进展。然而,实际系统通常同时使用摄像头和激光雷达,二者携带互补信息,而现有方法忽略了这一关键特性,导致生成结果仅覆盖单一模态(2D或3D)。为填补自动驾驶中2D-3D多模态联合生成的空白,本文提出框架HoloDrive,实现相机图像与激光雷达点云的联合生成。通过在异构生成模型间引入BEV-to-Camera与Camera-to-BEV转换模块,并在2D生成模型中加入深度预测分支以消除从图像空间反投影至鸟瞰空间的歧义;进一步扩展方法以支持时序预测,采用精心设计的渐进式训练策略。在单帧生成与世界建模基准上进行实验,结果表明,该方法在生成指标上显著优于当前最优(SOTA)方法。
原文摘要 · Abstract (English)
Generative models have significantly improved the generation and prediction quality on either camera images or LiDAR point clouds for autonomous driving. However, a real-world autonomous driving system uses multiple kinds of input modality, usually cameras and LiDARs, where they contain complementary information for generation, while existing generation methods ignore this crucial feature, resulting in the generated results only covering separate 2D or 3D information. In order to fill the gap in 2D-3D multi-modal joint generation for autonomous driving, in this paper, we propose our framework, \emph{HoloDrive}, to jointly generate the camera images and LiDAR point clouds. We employ BEV-to-Camera and Camera-to-BEV transform modules between heterogeneous generative models, and introduce a depth prediction branch in the 2D generative model to disambiguate the un-projecting from image space to BEV space, then extend the method to predict the future by adding temporal structure and carefully designed progressive training. Further, we conduct experiments on single frame generation and world model benchmarks, and demonstrate our method leads to significant performance gains over SOTA methods in terms of generation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。