生成逼真多视角驾驶视频与激光雷达序列,保持时空与模态一致性。
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
- 分两阶段融合扩散模型与3D-VAE,共享潜在空间实现视觉几何协同生成。
- 在nuScenes上达成FVD 16.95、FID 4.24、Chamfer 0.611,优于当前最佳。
- 适用于自动驾驶数据增强、感知任务训练,尤其适合需要多模态真实数据的场景。
我们提出Genesis,一个统一框架,用于联合生成多视角驾驶视频和激光雷达序列,同时保证时空一致性和跨模态一致性。Genesis采用两阶段架构,结合基于DiT的视频扩散模型与3D-VAE编码,并引入基于NeRF的渲染与自适应采样机制的BEV感知激光雷达生成器。两种模态通过共享潜在空间直接耦合,实现视觉与几何域的连贯演化。为提供结构化语义引导,我们设计DataCrafter,一个基于视觉语言模型的标注模块,可提供场景级与实例级监督。在nuScenes基准上的大量实验表明,Genesis在视频与激光雷达指标上均达到领先水平(FVD 16.95,FID 4.24,Chamfer 0.611),并显著提升下游任务如分割与3D检测性能,验证了生成数据的语义保真度与实际应用价值。
原文摘要 · Abstract (English)
We present Genesis, a unified framework for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-aware LiDAR generator with NeRF-based rendering and adaptive sampling. Both modalities are directly coupled through a shared latent space, enabling coherent evolution across visual and geometric domains. To guide the generation with structured semantics, we introduce DataCrafter, a captioning module built on vision-language models that provides scene-level and instance-level supervision. Extensive experiments on the nuScenes benchmark demonstrate that Genesis achieves state-of-the-art performance across video and LiDAR metrics (FVD 16.95, FID 4.24, Chamfer 0.611), and benefits downstream tasks including segmentation and 3D detection, validating the semantic fidelity and practical utility of the generated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。