arXiv:2412.03934cs.CVcs.AI2024-12ICCV被引 59

生成无限延伸且可控的高保真动态驾驶场景。

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

  • 用地图引导的稀疏体素模型生成无限空间3D世界。
  • 通过像素对齐引导缓冲区让视频模型生成外观一致的画面。
  • 支持通过地图、车辆框或文字控制动态物体,适合自动驾驶仿真。

我们提出InfiniCube,一种可扩展的高保真、可控制动态3D驾驶场景生成方法。现有方法在规模或几何/外观一致性上存在缺陷。我们利用最新的可扩展3D表示与视频模型,实现大规模动态场景生成,并支持通过高精地图、车辆边界框和文本描述进行灵活控制。首先,构建地图条件下的稀疏体素3D生成模型,以实现无界体素世界生成;其次,重新利用视频模型,并通过精心设计的像素对齐引导缓冲区将其锚定于体素世界,实现外观一致性合成;最后,提出一种快速前馈方法,结合体素与像素分支,将动态视频提升为可控制物体的动态3D高斯表示。大量实验验证了模型的有效性与优越性。

原文摘要 · Abstract (English)

We present InfiniCube, a scalable method for generating unbounded dynamic 3D driving scenes with high fidelity and controllability. Previous methods for scene generation either suffer from limited scales or lack geometric and appearance consistency along generated sequences. In contrast, we leverage the recent advancements in scalable 3D representation and video models to achieve large dynamic scene generation that allows flexible controls through HD maps, vehicle bounding boxes, and text descriptions. First, we construct a map-conditioned sparse-voxel-based 3D generative model to unleash its power for unbounded voxel world generation. Then, we re-purpose a video model and ground it on the voxel world through a set of carefully designed pixel-aligned guidance buffers, synthesizing a consistent appearance. Finally, we propose a fast feed-forward approach that employs both voxel and pixel branches to lift the dynamic videos to dynamic 3D Gaussians with controllable objects. Our method can generate controllable and realistic 3D driving scenes, and extensive experiments validate the effectiveness and superiority of our model.

3D生成自动驾驶视频建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。