任意场景生成:从任意俯视图生成可控驾驶视频
AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

- 用时空占据扩散变换器,从俯视图自回归生成占据序列
- 支持跨数据集布局和长时序生成,视频多视角一致
- 适合需要高可控合成数据的自动驾驶研究者
生成高质量且可控制的合成数据对端到端自动驾驶发展至关重要,尤其针对罕见安全关键场景。现有基于占据的方法通常依赖浅层条件机制和参考帧依赖的视频生成,限制了细粒度控制能力及可扩展性。本文提出AnyScene,一种统一的以占据为中心的驾驶场景生成框架。该框架通过时空占据扩散变换器,自回归地联合编码俯视图与占据特征,从任意俯视布局生成语义占据序列,实现跨数据集和用户定义布局的精准控制,并自然支持长时序生成。在此基础上,几何引导视角扩展模块将占据作为标准空间表示,以无参考、自回归方式合成时间一致的多视角驾驶视频,推理时支持灵活相机配置。大量实验表明,AnyScene在占据与视频生成上均达到领先性能,对未见及自定义布局具有强泛化能力,并为稀疏视角3D重建等下游任务带来显著提升。
原文摘要 · Abstract (English)
Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long tail of rare safety-critical scenarios. Existing occupancy-guided methods typically rely on shallow conditioning mechanisms and reference-frame-dependent video synthesis, which limits fine-grained controllability from arbitrary BEV layouts and restricts their applicability for scalable simulation. In this paper, we propose AnyScene, a unified occupancy-centric framework for driving scene generation. AnyScene generates semantic occupancy sequences from BEV layouts through a Spatial-Temporal Occupancy Diffusion Transformer that jointly tokenizes BEV and occupancy features in an autoregressive manner. This design enables precise controllability from cross-dataset and user-defined BEV inputs while naturally supporting long-horizon generation. Building upon the generated occupancy, a Geometry-Grounded View Expansion module treats occupancy as the canonical spatial representation and synthesizes temporally consistent multi-view driving videos in a reference-free and autoregressive fashion, supporting flexible camera configurations at inference time. Extensive experiments demonstrate that AnyScene achieves state-of-the-art performance in both occupancy and video generation. It exhibits strong generalization to unseen and customized layouts, and provides measurable benefits for downstream tasks such as sparse-view 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。