用流模型生成无限延伸3D世界,结构更准、速度更快。
WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation
- 基于流匹配思想,直接在3D数据分布间流动生成
- 无需隐空间,生成结构因果合理且收敛更快
- 支持向量控制布局与纹理,跨域生成效果好
无界3D世界生成正成为计算机视觉、图形学与机器人领域场景建模的基础任务。本文提出WorldFlow3D,一种可生成无界3D世界的新型方法。基于流匹配的核心特性——定义两个数据分布间的传输路径,我们将3D生成建模为在3D数据分布间流动的过程,突破了传统条件去噪的限制。所提无隐空间流方法生成的结构具有因果性和准确性,并可作为中间分布引导复杂结构与高质量纹理生成,且收敛速度优于现有方法。通过向量化的场景布局条件实现几何结构控制,利用场景属性实现视觉纹理控制。在真实户外驾驶场景与合成室内场景上均验证了方法的有效性,展现出良好的跨域泛化能力与真实数据分布上的高质量生成性能。在所有测试设置下,生成场景保真度均优于现有无界场景生成方法。
原文摘要 · Abstract (English)
Unbounded 3D world generation is emerging as a foundational task for scene modeling in computer vision, graphics, and robotics. In this work, we present WorldFlow3D, a novel method capable of generating unbounded 3D worlds. Building upon a foundational property of flow matching - namely, defining a path of transport between two data distributions - we model 3D generation more generally as a problem of flowing through 3D data distributions, not limited to conditional denoising. We find that our latent-free flow approach generates causal and accurate 3D structure, and can use this as an intermediate distribution to guide the generation of more complex structure and high-quality texture - all while converging more rapidly than existing methods. We enable controllability over generated scenes with vectorized scene layout conditions for geometric structure control and visual texture control through scene attributes. We confirm the effectiveness of WorldFlow3D on both real outdoor driving scenes and synthetic indoor scenes, validating cross-domain generalizability and high-quality generation on real data distributions. We confirm favorable scene generation fidelity over approaches in all tested settings for unbounded scene generation. For more, see https://light.princeton.edu/worldflow3d.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。