生成1024x576高分辨率驾驶视频,支持多视角3D控制。
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
- 基于双向调制变压器,精准对齐3D结构信息。
- 在nuScenes上达成FID 8.34、FVD 76.39的顶尖性能。
- 可低帧率条件下生成10Hz高清视频,适合自动驾驶训练。
生成模型的进展为合成逼真驾驶视频提供了有前景的解决方案,这对训练自动驾驶感知模型至关重要。然而,现有方法在多视角视频生成方面常因整合3D信息的挑战而表现不佳,难以维持时空一致性并从统一模型中有效学习。我们提出DriveScape,一个端到端的多视角、3D条件引导视频生成框架,可生成1024 x 576分辨率、10Hz的高质量视频。不同于受限于3D框标注帧率(通常仅2Hz)的其他方法,DriveScape在稀疏条件下仍能运行。其提出的双向调制变压器(BiMot)确保了3D结构信息的精确对齐,保持了时空一致性。在nuScenes数据集上,该方法达到最优性能,FID得分为8.34,FVD得分为76.39。
原文摘要 · Abstract (English)
Recent advancements in generative models have provided promising solutions for synthesizing realistic driving videos, which are crucial for training autonomous driving perception models. However, existing approaches often struggle with multi-view video generation due to the challenges of integrating 3D information while maintaining spatial-temporal consistency and effectively learning from a unified model. We propose DriveScape, an end-to-end framework for multi-view, 3D condition-guided video generation, capable of producing 1024 x 576 high-resolution videos at 10Hz. Unlike other methods limited to 2Hz due to the 3D box annotation frame rate, DriveScape overcomes this with its ability to operate under sparse conditions. Our Bi-Directional Modulated Transformer (BiMot) ensures precise alignment of 3D structural information, maintaining spatial-temporal consistency. DriveScape excels in video generation performance, achieving state-of-the-art results on the nuScenes dataset with an FID score of 8.34 and an FVD score of 76.39. Our project homepage: https://metadrivescape.github.io/papers_project/drivescapev1/index.html
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。