用轻量模块让多视角扩散模型直接生成可光照的3D资产
DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
- 通过模块化设计复用多视角扩散模型知识,实现几何与材质统一建模
- 仅用6.9万组多视角图像即可高效微调,生成高质量可光照3D网格资产
- 适合需要快速生成高保真3D内容的数字资产开发者
基于物理渲染(PBR)材质的3D资产创作耗时且依赖经验,亟需自主生成流程。现有方法多关注几何建模,或仅用顶点颜色烘焙纹理,或依赖后期图像扩散模型合成贴图。为实现端到端的PBR就绪3D资产生成,本文提出轻量级高斯资产适配器(LGAA),从新视角整合多视图(MV)扩散先验,统一建模几何与PBR材质。LGAA包含三个组件:LGAA Wrapper复用并适配多视图扩散模型网络层,利用其在数十亿图像上学习的知识,实现数据高效的快速收敛;LGAA Switcher对齐多个封装不同知识的Wrapper层,融合多扩散先验以同步生成几何与材质;设计一种受控变分自编码器(VAE),即LGAA Decoder,用于预测带PBR通道的2D高斯溅射(2DGS)。最后引入专用后处理流程,从生成的2DGS中有效提取高质量、可光照的网格资产。大量定量与定性实验表明,LGAA在文本和图像条件下的多视图扩散模型上均表现优越。模块化设计支持灵活集成多扩散先验,知识保留机制有效保持大规模图像数据集上学到的2D先验,仅需69,000个多视角实例即可完成微调,显著提升多视图扩散模型的3D生成能力。
原文摘要 · Abstract (English)
The labor- and experience-intensive creation of 3D assets with physically based rendering (PBR) materials demands an autonomous 3D asset creation pipeline. However, most existing 3D generation methods focus on geometry modeling, either baking textures into simple vertex colors or leaving texture synthesis to post-processing with image diffusion models. To achieve end-to-end PBR-ready 3D asset generation, we present Lightweight Gaussian Asset Adapter (LGAA), a novel framework that unifies the modeling of geometry and PBR materials by exploiting multi-view (MV) diffusion priors from a novel perspective. The LGAA features a modular design with three components. Specifically, the LGAA Wrapper reuses and adapts network layers from MV diffusion models, which encapsulate knowledge acquired from billions of images, enabling better convergence in a data-efficient manner. To incorporate multiple diffusion priors for geometry and PBR synthesis, the LGAA Switcher aligns multiple LGAA Wrapper layers encapsulating different knowledge. Then, a tamed variational autoencoder (VAE), termed LGAA Decoder, is designed to predict 2D Gaussian Splatting (2DGS) with PBR channels. Finally, we introduce a dedicated post-processing procedure to effectively extract high-quality, relightable mesh assets from the resulting 2DGS. Extensive quantitative and qualitative experiments demonstrate the superior performance of LGAA with both text- and image-conditioned MV diffusion models. Additionally, the modular design enables flexible incorporation of multiple diffusion priors, and the knowledge-preserving scheme effectively preseves the 2D priors learned on massive image dataset, which leads to data efficient finetuning to lift the MV diffuison models for 3D generation with merely 69k multi-view instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。