统一生成自动驾驶多模态传感器数据,解决真实数据难获取问题
OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
- 用统一的鸟瞰图空间融合多模态特征,实现跨传感器对齐
- 提出UAE方法,通过体渲染联合重建激光雷达与多视角图像
- 支持可控生成,可灵活调整传感器配置,适合仿真系统开发
自动驾驶发展依赖大量真实世界数据,但获取多样性和极端场景数据成本高、效率低。生成模型可通过合成逼真传感器数据提供解决方案,但现有方法多聚焦单模态生成,导致多模态数据不一致且效率低下。为此,我们提出OmniGen,一种在统一框架中生成对齐多模态传感器数据的新方法。该方法利用共享的鸟瞰图(BEV)空间统一多模态特征,并设计新型通用多模态重建方法UAE,通过体渲染联合解码激光雷达与多视角相机数据,实现精确灵活的重建。此外,引入带有ControlNet分支的扩散变压器(DiT),实现可控的多模态数据生成。全面实验表明,OmniGen在统一生成多模态传感器数据方面表现优异,具备良好的多模态一致性与灵活的传感器调节能力。
原文摘要 · Abstract (English)
Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising solution by synthesizing realistic sensor data. However, existing approaches primarily focus on single-modality generation, leading to inefficiencies and misalignment in multimodal sensor data. To address these challenges, we propose OminiGen, which generates aligned multimodal sensor data in a unified framework. Our approach leverages a shared Bird\u2019s Eye View (BEV) space to unify multimodal features and designs a novel generalizable multimodal reconstruction method, UAE, to jointly decode LiDAR and multi-view camera data. UAE achieves multimodal sensor decoding through volume rendering, enabling accurate and flexible reconstruction. Furthermore, we incorporate a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation. Our comprehensive experiments demonstrate that OminiGen achieves desired performances in unified multimodal sensor data generation with multimodal consistency and flexible sensor adjustments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。