用3D场景条件生成更一致的自动驾驶视频
CoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving
- 用精细3D表示替代2D布局,提升视频空间一致性
- 生成视频保持高几何保真度和视觉真实感
- 适合需要可控、多视角驾驶视频的自动驾驶研究
近期驾驶视频生成进展展现出增强自动驾驶系统训练数据可扩展性和可控性的巨大潜力。尽管基于预训练先进生成模型并由2D布局条件(如高清地图和边界框)引导的方法能生成逼真的驾驶视频,但实现具有高3D一致性的可控多视角视频仍是重大挑战。为此,我们提出一种新颖的空间自适应生成框架CoGen,利用3D生成技术在两个关键方面提升性能:(i) 为确保3D一致性,我们首先生成高质量、可控的3D条件,捕捉驾驶场景的几何结构。通过用细粒度3D表示替代粗略的2D条件,显著提升了生成视频的空间一致性。(ii) 此外,引入一致性适配模块以增强模型对多条件控制的鲁棒性。结果表明,该方法在保持几何保真度和视觉真实性方面表现优异,为自动驾驶提供了一种可靠的视频生成解决方案。
原文摘要 · Abstract (English)
Recent progress in driving video generation has shown significant potential for enhancing self-driving systems by providing scalable and controllable training data. Although pretrained state-of-the-art generation models, guided by 2D layout conditions (e.g., HD maps and bounding boxes), can produce photorealistic driving videos, achieving controllable multi-view videos with high 3D consistency remains a major challenge. To tackle this, we introduce a novel spatial adaptive generation framework, CoGen, which leverages advances in 3D generation to improve performance in two key aspects: (i) To ensure 3D consistency, we first generate high-quality, controllable 3D conditions that capture the geometry of driving scenes. By replacing coarse 2D conditions with these fine-grained 3D representations, our approach significantly enhances the spatial consistency of the generated videos. (ii) Additionally, we introduce a consistency adapter module to strengthen the robustness of the model to multi-condition control. The results demonstrate that this method excels in preserving geometric fidelity and visual realism, offering a reliable video generation solution for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。