Bolt3D仅需单张图像,7秒内生成高保真3D场景。
Bolt3D: Generating 3D Scenes in Seconds
- 基于2D扩散模型架构,直接生成3D场景表示。
- 单卡7秒完成生成,推理成本降低300倍。
- 适合快速原型设计与实时3D内容生成者。
我们提出一种用于快速前馈式3D场景生成的潜在扩散模型。给定一张或多张图像,Bolt3D在单个GPU上可在7秒内直接采样出3D场景表示。通过利用现有强大的可扩展2D扩散网络架构,实现一致且高保真的3D场景重建。为训练该模型,我们采用最先进的密集3D重建技术,对现有多视角图像数据集构建大规模多视角一致的3D几何与外观数据集。相较于以往需要针对每个场景进行优化的多视角生成模型,Bolt3D将推理成本降低至原来的1/300。
原文摘要 · Abstract (English)
We present a latent diffusion model for fast feed-forward 3D scene generation. Given one or more images, our model Bolt3D directly samples a 3D scene representation in less than seven seconds on a single GPU. We achieve this by leveraging powerful and scalable existing 2D diffusion network architectures to produce consistent high-fidelity 3D scene representations. To train this model, we create a large-scale multiview-consistent dataset of 3D geometry and appearance by applying state-of-the-art dense 3D reconstruction techniques to existing multiview image datasets. Compared to prior multiview generative models that require per-scene optimization for 3D reconstruction, Bolt3D reduces the inference cost by a factor of up to 300 times.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。