arXiv:2501.13718cs.CV2025-01

用互信息分析生成模型中隐变量的作用,提升合成数据质量。

A Mutual Information Perspective on Multiple Latent Variable Generative Models for Positive View Generation

  • 用互信息量化每个隐变量的贡献,揭示现有模型利用不充分
  • 基于层级解耦隐变量生成多样且语义明确的合成视图
  • 动态采样策略在自监督学习中显著提升数据多样性

在图像生成中,多隐变量生成模型(MLVGMs)通过多个隐变量逐步构建图像,从全局特征到局部细节(如StyleGAN、NVAE),已成为广泛应用的强大工具。然而其生成机制仍仅靠经验观察,缺乏对各隐变量作用的系统理解。本文提出一种新框架,以互信息(MI)为度量,量化每个隐变量的贡献。分析发现当前MLVGM常未充分利用部分隐变量,并据此提出在下游应用中的改进思路。在此基础上,我们设计一种用于自监督对比表征学习(SSCRL)的合成数据生成方法。利用MLVGM的层级与解耦特性,无需真实图像即可生成多样且语义合理的视图。此外,引入连续采样(CS)策略,在SSCRL训练过程中动态生成新样本,大幅提升数据变异性。大量实验表明,所生成视图性能可媲美甚至超越真实数据生成的视图。该工作建立了一种原则性方法,深入理解并有效利用MLVGM,推动生成建模与自监督学习的发展。代码与预训练模型见:https://github.com/SerezD/mi_ml_gen。

原文摘要 · Abstract (English)

In image generation, Multiple Latent Variable Generative Models (MLVGMs) employ multiple latent variables to gradually shape the final images, from global characteristics to finer and local details (e.g., StyleGAN, NVAE), emerging as powerful tools for diverse applications. Yet their generative dynamics remain only empirically observed, without a systematic understanding of each latent variable's impact. In this work, we propose a novel framework that quantifies the contribution of each latent variable using Mutual Information (MI) as a metric. Our analysis reveals that current MLVGMs often underutilize some latent variables, and provides actionable insights for their use in downstream applications. With this foundation, we introduce a method for generating synthetic data for Self-Supervised Contrastive Representation Learning (SSCRL). By leveraging the hierarchical and disentangled variables of MLVGMs, our approach produces diverse and semantically meaningful views without the need for real image data. Additionally, we introduce a Continuous Sampling (CS) strategy, where the generator dynamically creates new samples during SSCRL training, greatly increasing data variability. Our comprehensive experiments demonstrate the effectiveness of these contributions, showing that MLVGMs' generated views compete on par with or even surpass views generated from real data. This work establishes a principled approach to understanding and exploiting MLVGMs, advancing both generative modeling and self-supervised learning. Code and pre-trained models at: https://github.com/SerezD/mi_ml_gen.

生成模型互信息自监督学习合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。