arXiv:2505.17590cs.CV2025-05NeurIPS被引 7

无需视角条件控制,实现高分辨率人脸3D一致生成

CGS-GAN: 3D Consistent Gaussian Splatting GANs for High Resolution Human Head Synthesis

  • 不依赖视角条件,通过多视图正则化稳定训练
  • 支持2048²高分辨率输出,FID表现优异
  • 适合需要高质量3D人脸生成的研究与应用

近期基于3D高斯点阵的3D GAN已用于高质量人脸合成,但现有方法通过将随机隐变量与当前相机位置关联来稳定训练并提升视角质量,导致重渲染时身份随视角变化明显,破坏3D一致性。固定视角虽可得高质量单一视角,但新视角效果差。去除视角条件常致训练崩溃。为此,我们提出CGS-GAN,一种无需视角条件即可实现稳定训练与高质量3D一致人脸合成的新框架。引入多视图正则化提升生成器收敛性,计算开销极小;改进条件损失,并设计适配架构,既稳定训练又支持高效渲染与扩展,最高输出分辨率达$2048^2$。我们构建新数据集(基于FFHQ),聚焦更大头部区域,减少视角相关伪影,剔除手部遮挡等图像,显著提升3D一致性。实验表明,该方法在高分辨率下保持优质渲染,FID分数具有竞争力,确保3D场景生成一致性。

原文摘要 · Abstract (English)

Recently, 3D GANs based on 3D Gaussian splatting have been proposed for high quality synthesis of human heads. However, existing methods stabilize training and enhance rendering quality from steep viewpoints by conditioning the random latent vector on the current camera position. This compromises 3D consistency, as we observe significant identity changes when re-synthesizing the 3D head with each camera shift. Conversely, fixing the camera to a single viewpoint yields high-quality renderings for that perspective but results in poor performance for novel views. Removing view-conditioning typically destabilizes GAN training, often causing the training to collapse. In response to these challenges, we introduce CGS-GAN, a novel 3D Gaussian Splatting GAN framework that enables stable training and high-quality 3D-consistent synthesis of human heads without relying on view-conditioning. To ensure training stability, we introduce a multi-view regularization technique that enhances generator convergence with minimal computational overhead. Additionally, we adapt the conditional loss used in existing 3D Gaussian splatting GANs and propose a generator architecture designed to not only stabilize training but also facilitate efficient rendering and straightforward scaling, enabling output resolutions up to $2048^2$. To evaluate the capabilities of CGS-GAN, we curate a new dataset derived from FFHQ. This dataset enables very high resolutions, focuses on larger portions of the human head, reduces view-dependent artifacts for improved 3D consistency, and excludes images where subjects are obscured by hands or other objects. As a result, our approach achieves very high rendering quality, supported by competitive FID scores, while ensuring consistent 3D scene generation. Check our our project page here: https://fraunhoferhhi.github.io/cgs-gan/

3D生成高斯点阵人脸合成图像质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。