一秒生成高质量3D场景,速度比之前快10到100倍。
FlashWorld: High-quality 3D Scene Generation within Seconds
- 直接生成3D高斯表示,跳过传统多视角重建流程。
- 在保持3D一致性前提下,视觉质量显著提升,推理步数减少。
- 适合需要快速生成高保真3D内容的创作者和开发者。
我们提出FlashWorld,一种从单张图像或文本提示中在数秒内生成高质量3D场景的生成模型,速度比以往工作快10~100倍,同时渲染质量更优。该方法摒弃传统的多视角导向(MV-oriented)范式,转而采用3D导向策略,在生成多视角图像时直接输出3D高斯表示。尽管3D导向方法通常视觉质量较差,但FlashWorld通过双模式预训练与跨模式后训练实现优势互补:首先基于视频扩散模型预训练一个支持多视角与3D导向两种模式的多视图扩散模型;随后引入跨模式蒸馏,将3D导向模式的分布对齐至高质量多视角模式,从而在保持3D一致性的同时显著提升视觉质量,并减少推理所需去噪步数。此外,我们设计策略利用大量单视角图像与文本提示增强模型对分布外输入的泛化能力。大量实验验证了该方法在效率与性能上的优越性。
原文摘要 · Abstract (English)
We propose FlashWorld, a generative model that produces 3D scenes from a single image or text prompt in seconds, 10~100$\times$ faster than previous works while possessing superior rendering quality. Our approach shifts from the conventional multi-view-oriented (MV-oriented) paradigm, which generates multi-view images for subsequent 3D reconstruction, to a 3D-oriented approach where the model directly produces 3D Gaussian representations during multi-view generation. While ensuring 3D consistency, 3D-oriented method typically suffers poor visual quality. FlashWorld includes a dual-mode pre-training phase followed by a cross-mode post-training phase, effectively integrating the strengths of both paradigms. Specifically, leveraging the prior from a video diffusion model, we first pre-train a dual-mode multi-view diffusion model, which jointly supports MV-oriented and 3D-oriented generation modes. To bridge the quality gap in 3D-oriented generation, we further propose a cross-mode post-training distillation by matching distribution from consistent 3D-oriented mode to high-quality MV-oriented mode. This not only enhances visual quality while maintaining 3D consistency, but also reduces the required denoising steps for inference. Also, we propose a strategy to leverage massive single-view images and text prompts during this process to enhance the model's generalization to out-of-distribution inputs. Extensive experiments demonstrate the superiority and efficiency of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。