用空间填充向量量化揭示GAN隐空间的可解释结构
Unsupervised Panoptic Interpretation of Latent Spaces in GANs Using Space-Filling Vector Quantization
- 提出空间填充向量量化(SFVQ)方法,沿分段线性曲线对隐空间进行量化
- 在StyleGAN2和BigGAN上验证,SFVQ曲线能定位生成因子对应的区域
- 可用于可控数据增强和可解释图像变换,适合研究生成模型可解释性者
生成对抗网络(GAN)学习的隐空间可映射为真实图像,但其结构难以解释。以往监督方法需依赖标签或标注样本训练,而本文提出一种向量量化改进方法——空间填充向量量化(SFVQ),将数据量化到分段线性曲线上,能够捕捉隐空间的潜在形态结构,从而实现可解释性。我们将其应用于预训练的StyleGAN2与BigGAN模型,在多个数据集上验证:SFVQ曲线能普遍建立隐空间的可解释模型,明确特定区域对应生成因子;且曲线上的每条线可代表一个可解释方向,用于直观图像变换。此外,位于SFVQ曲线上的点可用于可控数据增强。
原文摘要 · Abstract (English)
Generative adversarial networks (GANs) learn a latent space whose samples can be mapped to real-world images. Such latent spaces are difficult to interpret. Some earlier supervised methods aim to create an interpretable latent space or discover interpretable directions, which requires exploiting data labels or annotated synthesized samples for training. However, we propose using a modification of vector quantization called space-filling vector quantization (SFVQ), which quantizes the data on a piece-wise linear curve. SFVQ can capture the underlying morphological structure of the latent space, making it interpretable. We apply this technique to model the latent space of pre-trained StyleGAN2 and BigGAN networks on various datasets. Our experiments show that the SFVQ curve yields a general interpretable model of the latent space such that it determines which parts of the latent space correspond to specific generative factors. Furthermore, we demonstrate that each line of the SFVQ curve can potentially refer to an interpretable direction for applying intelligible image transformations. We also demonstrate that the points located on an SFVQ line can be used for controllable data augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。