arXiv:2506.19117cs.CV2025-06中稿 · CVPR

用可编辑的3D组件生成城市场景,更省内存、更快、更好调。

PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes

  • 用向量化的物体和栅格化地面构建混合表示,支持结构化操作
  • 在KITTI-360上实现最优生成质量,内存更低、推理更快
  • 适合需要可控编辑的城市场景生成任务

现有3D语义城市场景生成方法多依赖体素表示,存在分辨率固定、难编辑、内存占用高等问题。本文提出基于几何原型的范式,用紧凑且语义明确的3D元素构建场景,更易操作与组合。为此,我们引入PrITTI——一种利用向量化物体原型和栅格化地面的潜在扩散模型,生成多样化、可控制、可编辑的3D语义城市场景。该混合表示构建出结构化的潜在空间,支持物体级与地面级操作。在KITTI-360上的实验表明,基于原型的表示充分发挥了扩散变换器的能力,相比体素方法,在生成质量上达到当前最优,同时内存更低、推理更快、编辑性更强。除生成外,PrITTI还支持场景编辑、补全、扩展及逼真街景合成等下游应用。代码与更多结果见https://raniatze.github.io/pritti/。

原文摘要 · Abstract (English)

Existing approaches to 3D semantic urban scene generation predominantly rely on voxel-based representations, which are bound by fixed resolution, challenging to edit, and memory-intensive in their dense form. In contrast, we advocate for a primitive-based paradigm where urban scenes are represented using compact, semantically meaningful 3D elements that are easy to manipulate and compose. To this end, we introduce PrITTI, a latent diffusion model that leverages vectorized object primitives and rasterized ground surfaces for generating diverse, controllable, and editable 3D semantic urban scenes. This hybrid representation yields a structured latent space that facilitates object- and ground-level manipulation. Experiments on KITTI-360 show that primitive-based representations unlock the full capabilities of diffusion transformers, achieving state-of-the-art 3D scene generation quality with lower memory requirements, faster inference, and greater editability than voxel-based methods. Beyond generation, PrITTI supports a range of downstream applications, including scene editing, inpainting, outpainting, and photo-realistic street-view synthesis. The source code and more results can be found at https://raniatze.github.io/pritti/.

3D生成城市场景扩散模型可编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。