用单图生成大型3D场景,还能精细控制布局。
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

- 把图像转3D模型技术扩展为场景级生成,通过卷积化改造模型。
- 在自研合成数据上微调,生成任意大小复杂度的3D场景。
- 适合需要高保真3D场景生成的研究者和创作者。
我们提出SynCity 3000,一种生成全局一致且可精细控制布局的3D场景的框架。基于当前图像到3D生成器能从单张图像生成复杂3D资产的能力,我们将该能力扩展至整个场景规模,通过将生成器改造成卷积算子实现。为此,我们设计了一种新的合成数据引擎,以解决3D场景训练数据稀缺的问题,并在生成的场景数据上对模型进行微调。随后,将该卷积生成器应用于由用户提示生成的正交视图图像,从而输出任意大小与复杂度的3D场景。在多种提示和布局下,SynCity 3000均能生成大尺度、连贯且细节丰富的3D场景,有效弥补了先前3D场景生成方法的不足。
原文摘要 · Abstract (English)
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。