arXiv:2503.07152cs.CV2025-03ICCV被引 11

用场景图控制生成高保真3D户外场景,支持交互式修改。

Controllable 3D Outdoor Scene Generation via Scene Graphs

  • 通过场景图生成密集鸟瞰图嵌入图,引导扩散模型生成3D场景。
  • 在真实城市场景数据上实现与输入图高度一致的高质量生成结果。
  • 首个基于场景图生成3D户外场景的方法,适合游戏/自动驾驶场景设计。

三维场景生成在计算机视觉中至关重要,广泛应用于自动驾驶、游戏和元宇宙等领域。现有方法要么缺乏用户控制,要么依赖不直观的条件。本文提出一种新方法,利用易于理解的场景图作为控制输入,生成户外3D场景。我们开发了一个交互系统,将稀疏场景图转换为密集的鸟瞰图(BEV)嵌入图,指导条件扩散模型生成与场景图描述一致的3D场景。推理阶段,用户可轻松创建或修改场景图以生成大规模户外场景。我们构建了一个包含配对场景图与3D语义场景的大规模数据集,用于训练BEV嵌入和扩散模型。实验表明,该方法能持续生成高质量的城市3D场景,且与输入场景图高度一致。据我们所知,这是首个基于场景图生成3D户外场景的方法。

原文摘要 · Abstract (English)

Three-dimensional scene generation is crucial in computer vision, with applications spanning autonomous driving, gaming and the metaverse. Current methods either lack user control or rely on imprecise, non-intuitive conditions. In this work, we propose a method that uses, scene graphs, an accessible, user friendly control format to generate outdoor 3D scenes. We develop an interactive system that transforms a sparse scene graph into a dense BEV (Bird's Eye View) Embedding Map, which guides a conditional diffusion model to generate 3D scenes that match the scene graph description. During inference, users can easily create or modify scene graphs to generate large-scale outdoor scenes. We create a large-scale dataset with paired scene graphs and 3D semantic scenes to train the BEV embedding and diffusion models. Experimental results show that our approach consistently produces high-quality 3D urban scenes closely aligned with the input scene graphs. To the best of our knowledge, this is the first approach to generate 3D outdoor scenes conditioned on scene graphs.

3D生成场景图扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。