用分层潜在树生成大尺度高质量3D场景,支持任意大小合成
LT3SD: Latent Trees for 3D Scene Diffusion
- 采用分层潜在树编码几何与细节,实现从粗到精的结构建模
- 在场景块上训练,通过共享扩散过程生成任意尺寸的大场景
- 适用于大规模无条件生成和部分观测下的概率补全任务
我们提出LT3SD,一种新型潜空间扩散模型,用于大规模3D场景生成。尽管扩散模型在3D物体生成上表现优异,但扩展至3D场景时存在空间范围与质量局限。为生成复杂多样的3D场景结构,我们引入潜空间树表示,以层次化方式有效编码低频几何与高频细节。在此潜空间中,我们学习生成式扩散过程,对每一分辨率层级的场景潜变量进行建模。为生成具有不同尺寸的大规模场景,我们在场景块上训练模型,并通过多个场景块间的共享扩散生成任意尺寸的输出3D场景。大量实验表明,LT3SD在大规模、高质量的无条件3D场景生成以及部分场景观测的概率补全任务中均表现出色。
原文摘要 · Abstract (English)
We present LT3SD, a novel latent diffusion model for large-scale 3D scene generation. Recent advances in diffusion models have shown impressive results in 3D object generation, but are limited in spatial extent and quality when extended to 3D scenes. To generate complex and diverse 3D scene structures, we introduce a latent tree representation to effectively encode both lower-frequency geometry and higher-frequency detail in a coarse-to-fine hierarchy. We can then learn a generative diffusion process in this latent 3D scene space, modeling the latent components of a scene at each resolution level. To synthesize large-scale scenes with varying sizes, we train our diffusion model on scene patches and synthesize arbitrary-sized output 3D scenes through shared diffusion generation across multiple scene patches. Through extensive experiments, we demonstrate the efficacy and benefits of LT3SD for large-scale, high-quality unconditional 3D scene generation and for probabilistic completion for partial scene observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。