X-Scene可生成高保真大场景驾驶环境,支持灵活控制与时空一致性。
X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability
- 分粒度控制:支持布局输入与语义提示双重调节。
- 多模态统一生成:依次产出语义体素与多视角图像视频。
- 扩展生成能力:通过一致性外推实现大范围场景无缝拓展。
扩散模型正推动自动驾驶发展,用于真实数据合成、端到端规划与闭环仿真,主要聚焦时序一致性生成。然而,需空间一致性的大规模3D场景生成仍研究不足。本文提出X-Scene框架,实现几何精细、外观逼真且灵活可控的大规模驾驶场景生成。具体而言,该框架支持多粒度控制:低层级由用户输入或文本驱动布局条件,实现细节场景构建;高层级则基于用户意图与大语言模型增强提示进行语义引导,提升定制效率。为增强几何与视觉保真度,提出统一流水线,顺序生成3D语义占用栅格及对应多视图图像与视频,确保跨模态对齐与时间一致性。进一步通过一致性感知外推,将局部区域扩展为大场景,从已有生成区域延伸体素与图像以维持空间与视觉连贯性。最终场景升维至高质量3DGS表示,支持仿真与场景探索等多种应用。大量实验表明,X-Scene显著提升大规模场景生成的可控性与保真度,赋能自动驾驶的数据生成与仿真。
原文摘要 · Abstract (English)
Diffusion models are advancing autonomous driving by enabling realistic data synthesis, predictive end-to-end planning, and closed-loop simulation, with a primary focus on temporally consistent generation. However, large-scale 3D scene generation requiring spatial coherence remains underexplored. In this paper, we present X-Scene, a novel framework for large-scale driving scene generation that achieves geometric intricacy, appearance fidelity, and flexible controllability. Specifically, X-Scene supports multi-granular control, including low-level layout conditioning driven by user input or text for detailed scene composition, and high-level semantic guidance informed by user intent and LLM-enriched prompts for efficient customization. To enhance geometric and visual fidelity, we introduce a unified pipeline that sequentially generates 3D semantic occupancy and corresponding multi-view images and videos, ensuring alignment and temporal consistency across modalities. We further extend local regions into large-scale scenes via consistency-aware outpainting, which extrapolates occupancy and images from previously generated areas to maintain spatial and visual coherence. The resulting scenes are lifted into high-quality 3DGS representations, supporting diverse applications such as simulation and scene exploration. Extensive experiments demonstrate that X-Scene substantially advances controllability and fidelity in large-scale scene generation, empowering data generation and simulation for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。