用扩散模型生成4.4万张3D人脸,实现语义可控的高质量人脸资产创建
Bringing Diversity from Diffusion Models to Semantic-Guided Face Asset Generation
- 利用预训练扩散模型生成高质3D人脸数据,再通过归一化模块转换为真实扫描风格
- 构建基于GAN的生成器,支持44,000张人脸数据训练,可按语义属性生成几何与材质
- 支持潜空间连续编辑,适合影视、游戏等需多样化人脸资产的场景
数字人脸建模广泛应用于多个领域,但受限于采集设备、人工成本和演员条件,难以获得多样性和可控性兼备的模型。本文提出一种语义可控的生成网络,通过引入基于预训练扩散模型的新数据生成流程,构建高质量3D人脸数据库。我们设计的归一化模块将扩散模型合成的数据转化为接近真实扫描的质量。基于生成的44,000张人脸模型,进一步开发了高效的GAN生成器,可接收语义属性输入,生成几何与反照率信息,并支持潜空间中的连续属性编辑。后续的资产精修组件生成物理真实的人脸资产。本文构建了一个完整的可交互系统,并集成至网页工具中,计划随论文公开。
原文摘要 · Abstract (English)
Digital modeling and reconstruction of human faces serve various applications. However, its availability is often hindered by the requirements of data capturing devices, manual labor, and suitable actors. This situation restricts the diversity, expressiveness, and control over the resulting models. This work aims to demonstrate that a semantically controllable generative network can provide enhanced control over the digital face modeling process. To enhance diversity beyond the limited human faces scanned in a controlled setting, we introduce a novel data generation pipeline that creates a high-quality 3D face database using a pre-trained diffusion model. Our proposed normalization module converts synthesized data from the diffusion model into high-quality scanned data. Using the 44,000 face models we obtained, we further developed an efficient GAN-based generator. This generator accepts semantic attributes as input, and generates geometry and albedo. It also allows continuous post-editing of attributes in the latent space. Our asset refinement component subsequently creates physically-based facial assets. We introduce a comprehensive system designed for creating and editing high-quality face assets. Our proposed model has undergone extensive experiment, comparison and evaluation. We also integrate everything into a web-based interactive tool. We aim to make this tool publicly available with the release of the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。