用可学习的高斯嵌入提升数字人生成的3D一致性和动作准确性
Controlling Avatar Diffusion with Learnable Gaussian Embedding
- 在参数化头部表面嵌入可优化的神经高斯作为控制信号
- 合成大规模多姿态多身份数据集,真实/合成标签区分提升生成质量
- 相比现有方法,在真实感、表达力和3D一致性上均有显著提升
扩散模型在数字人生成方面取得进展,但现有方法仍难以保持3D一致性、时间连贯性和动作准确性。主要原因是常用控制信号(如关键点、深度图)表达能力有限,且公开数据集在身份和姿态多样性上不足。本文分析了当前控制信号的缺陷,提出一种可优化、密集、表达性强且3D一致的新控制信号表示:将可学习的神经高斯嵌入到参数化头部表面,显著提升了基于扩散模型的头部生成的一致性与表现力。同时,构建大规模合成数据集,包含多种姿态与身份,并使用真实/合成标签有效区分真实与合成数据,降低合成数据缺陷对生成结果的影响。大量实验表明,该方法在真实感、表达力和3D一致性上均优于现有方法。代码、合成数据集及预训练模型将在项目主页发布:https://ustc3dv.github.io/Learn2Control/
原文摘要 · Abstract (English)
Recent advances in diffusion models have made significant progress in digital human generation. However, most existing models still struggle to maintain 3D consistency, temporal coherence, and motion accuracy. A key reason for these shortcomings is the limited representation ability of commonly used control signals(e.g., landmarks, depth maps, etc.). In addition, the lack of diversity in identity and pose variations in public datasets further hinders progress in this area. In this paper, we analyze the shortcomings of current control signals and introduce a novel control signal representation that is optimizable, dense, expressive, and 3D consistent. Our method embeds a learnable neural Gaussian onto a parametric head surface, which greatly enhances the consistency and expressiveness of diffusion-based head models. Regarding the dataset, we synthesize a large-scale dataset with multiple poses and identities. In addition, we use real/synthetic labels to effectively distinguish real and synthetic data, minimizing the impact of imperfections in synthetic data on the generated head images. Extensive experiments show that our model outperforms existing methods in terms of realism, expressiveness, and 3D consistency. Our code, synthetic datasets, and pre-trained models will be released in our project page: https://ustc3dv.github.io/Learn2Control/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。