用单帧生成无限延伸的动态城市场景,支持可控、连贯的自动驾驶训练环境。
InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

- 基于输入多视角图像重建3D占据表示,驱动沿任意轨迹扩展场景
- 生成视频FID 6.4,FVD 67.97,时长与稳定性显著优于现有方法
- 视觉与空间跨模态协同优化,适合自动驾驶仿真与数据增强
在自动驾驶领域,生成真实、可控且时空连贯的城市环境是一项关键但尚未解决的挑战。本文提出InfiniVerse,一个统一的流水线,可从单帧图像实现长距离、2D-3D对齐、可控的动态城市场景合成。首先,模型从多视角输入帧重建3D占据表示,作为沿任意轨迹自回归扩展场景的基础。随后,视频扩散模型将粗粒度占据网格转化为逼真、时空一致的视频序列。此外,我们提出分层草图-精修范式,将生成视频重投影为图像条件反馈,增强3D占据表示,实现视觉与空间域间的跨模态对齐与相互增强。在Waymo Open Dataset和nuScenes上的大量评估表明,InfiniVerse达到当前最佳性能,FID为6.4,FVD为67.97,在持续时间与稳定性上显著超越现有基准。
原文摘要 · Abstract (English)
Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。