用扩散模型一键生成3D资产参数,省时又精准。
DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation
- 用轻量扩散变换器直接预测3D生成参数
- 仅7.6M参数,30小时训练即实现高精度还原
- 适合游戏/影视行业快速生成高质量3D模型
程序化内容生成(PCG)在创建高质量3D内容方面具有强大能力,但控制其生成期望形状仍困难,通常需大量参数调优。逆向程序化内容生成旨在自动寻找满足输入条件的最佳参数。然而,现有基于采样的方法和神经网络方法仍存在迭代次数多或可控性差的问题。本文提出DI-PCG,一种从通用图像条件出发的高效逆向PCG新方法。其核心是一个轻量级扩散变换器模型,将PCG参数直接作为去噪目标,观测图像作为控制参数生成的条件。DI-PCG效率高、效果好:仅需7.6M网络参数和30 GPU小时训练,即可实现参数的高精度恢复,并在真实图像上表现出良好泛化能力。定量与定性实验验证了其在逆向PCG及图像到3D生成任务中的有效性。该方法为高效逆向PCG提供了可行路径,是基于参数模型构建3D资产生成流程的重要探索。
原文摘要 · Abstract (English)
Procedural Content Generation (PCG) is powerful in creating high-quality 3D contents, yet controlling it to produce desired shapes is difficult and often requires extensive parameter tuning. Inverse Procedural Content Generation aims to automatically find the best parameters under the input condition. However, existing sampling-based and neural network-based methods still suffer from numerous sample iterations or limited controllability. In this work, we present DI-PCG, a novel and efficient method for Inverse PCG from general image conditions. At its core is a lightweight diffusion transformer model, where PCG parameters are directly treated as the denoising target and the observed images as conditions to control parameter generation. DI-PCG is efficient and effective. With only 7.6M network parameters and 30 GPU hours to train, it demonstrates superior performance in recovering parameters accurately, and generalizing well to in-the-wild images. Quantitative and qualitative experiment results validate the effectiveness of DI-PCG in inverse PCG and image-to-3D generation tasks. DI-PCG offers a promising approach for efficient inverse PCG and represents a valuable exploration step towards a 3D generation path that models how to construct a 3D asset using parametric models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。