arXiv:2602.07198cs.CVcs.GR2026-02中稿 · ICLR被引 3

用语义特征替代视角做条件,让3D人脸生成更稳定多样。

Condition Matters in Full-head 3D GANs

  • 用正面图像的语义特征作为跨视角共享条件,消除视角偏差。
  • 在多个数据集上生成质量、多样性显著提升,尤其在非条件视角下表现更好。
  • 适合需要高质量3D人脸生成与泛化能力的研究者使用。

条件输入对全头3D GAN的稳定训练至关重要。以往方法仅以视角作为条件,导致生成空间沿视角方向产生偏差,造成不同视角间生成质量与多样性差异大,整体不一致。本文提出使用视图不变的语义特征作为条件,解耦生成能力与视角依赖。为此,我们利用FLUX.1 Kontext将现有高质量正面人脸数据集扩展至多视角,提取正面图像的clip特征作为所有视角的共享语义条件,实现语义对齐并消除方向偏差。该共享条件使同一主体的多视角监督得以整合,加速训练并增强生成头像的全局一致性。此外,语义条件有助于引导生成器持续学习真实语义分布,避免多样性停滞。大量实验表明,本方法在全头合成与单视图GAN反演任务中均显著提升了保真度、多样性和泛化能力。

原文摘要 · Abstract (English)

Conditioning is crucial for stable training of full-head 3D GANs. Without any conditioning signal, the model suffers from severe mode collapse, making it impractical to training. However, a series of previous full-head 3D GANs conventionally choose the view angle as the conditioning input, which leads to a bias in the learned 3D full-head space along the conditional view direction. This is evident in the significant differences in generation quality and diversity between the conditional view and non-conditional views of the generated 3D heads, resulting in global incoherence across different head regions. In this work, we propose to use view-invariant semantic feature as the conditioning input, thereby decoupling the generative capability of 3D heads from the viewing direction. To construct a view-invariant semantic condition for each training image, we create a novel synthesized head image dataset. We leverage FLUX.1 Kontext to extend existing high-quality frontal face datasets to a wide range of view angles. The image clip feature extracted from the frontal view is then used as a shared semantic condition across all views in the extended images, ensuring semantic alignment while eliminating directional bias. This also allows supervision from different views of the same subject to be consolidated under a shared semantic condition, which accelerates training and enhances the global coherence of the generated 3D heads. Moreover, as GANs often experience slower improvements in diversity once the generator learns a few modes that successfully fool the discriminator, our semantic conditioning encourages the generator to follow the true semantic distribution, thereby promoting continuous learning and diverse generation. Extensive experiments on full-head synthesis and single-view GAN inversion demonstrate that our method achieves significantly higher fidelity, diversity, and generalizability.

3D生成GAN语义条件人脸建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。