arXiv:2504.17712cs.CV2025-04

提出生成场理论,实现对StyleGAN特征合成的解耦控制。

Generative Fields: Uncovering Hierarchical Feature Control for StyleGAN via Inverted Receptive Fields

  • 基于卷积神经网络感受野思想,构建生成场理论解释层级特征合成。
  • 引入通道级风格潜空间S,直接在生成时解耦控制图像特征。
  • 无需预训练即可编辑,适用于需要精细控制的图像生成场景。

StyleGAN能从随机噪声生成高度逼真的虚拟人脸,但其低维潜在空间中特征强烈纠缠,难以控制。以往工作通过调节更具表现力的W空间采样来实现图像或文本提示控制,但W空间仍无法直接控制特征生成,且其中的特征嵌入需预训练重建风格信号,限制了应用。本文提出“生成场”概念,借鉴卷积神经网络的感受野机制,解释StyleGAN中的层级特征合成过程。进一步提出基于生成场理论与通道级风格潜空间S的新图像编辑流程,利用CNN内在结构,在生成时实现对特征合成的解耦控制,无需预训练即可直接操作。

原文摘要 · Abstract (English)

StyleGAN has demonstrated the ability of GANs to synthesize highly-realistic faces of imaginary people from random noise. One limitation of GAN-based image generation is the difficulty of controlling the features of the generated image, due to the strong entanglement of the low-dimensional latent space. Previous work that aimed to control StyleGAN with image or text prompts modulated sampling in W latent space, which is more expressive than Z latent space. However, W space still has restricted expressivity since it does not control the feature synthesis directly; also the feature embedding in W space requires a pre-training process to reconstruct the style signal, limiting its application. This paper introduces the concept of "generative fields" to explain the hierarchical feature synthesis in StyleGAN, inspired by the receptive fields of convolution neural networks (CNNs). Additionally, we propose a new image editing pipeline for StyleGAN using generative field theory and the channel-wise style latent space S, utilizing the intrinsic structural feature of CNNs to achieve disentangled control of feature synthesis at synthesis time.

生成模型图像编辑风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。