用高效扩散模型生成可控制的长时3D驾驶场景,兼顾质量与速度。
DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion
- 结合文本和俯视图引导,用轻量Transformer加速视频生成。
- 支持800×424分辨率、12帧/秒的8秒长视频与3D重建。
- 适合自动驾驶仿真与高精度场景生成需求者使用。
我们提出DriveGen3D,一种新型框架,用于生成高质量且高度可控的动态3D驾驶场景,解决现有方法在长期生成效率、3D表示缺失或仅限静态重建等方面的局限。该框架通过多模态条件控制,将高效长时视频生成与大规模动态场景重建统一起来。其包含两个核心模块:FastDrive-DiT,一种高效的视频扩散Transformer,可在文本与鸟瞰图(BEV)布局引导下生成高分辨率、时间连贯的视频;FastRecon3D,一个前馈模块,可快速构建随时间变化的3D高斯表示,保证时空一致性。该系统能生成长达8秒(800×424,12 FPS)的驾驶视频及对应3D场景,达到当前最优性能的同时保持高效。
原文摘要 · Abstract (English)
We present DriveGen3D, a novel framework for generating high-quality and highly controllable dynamic 3D driving scenes that addresses critical limitations in existing methodologies. Current approaches to driving scene synthesis either suffer from prohibitive computational demands for extended temporal generation, focus exclusively on prolonged video synthesis without 3D representation, or restrict themselves to static single-scene reconstruction. Our work bridges this methodological gap by integrating accelerated long-term video generation with large-scale dynamic scene reconstruction through multimodal conditional control. DriveGen3D introduces a unified pipeline consisting of two specialized components: FastDrive-DiT, an efficient video diffusion transformer for high-resolution, temporally coherent video synthesis under text and Bird's-Eye-View (BEV) layout guidance; and FastRecon3D, a feed-forward module that rapidly builds 3D Gaussian representations across time, ensuring spatial-temporal consistency. DriveGen3D enable the generation of long driving videos (up to $800\times424$ at $12$ FPS) and corresponding 3D scenes, achieving state-of-the-art results while maintaining efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。