看似独立的风格变量其实仍泄露类别信息,仅匹配边缘分布不够。
Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models

- 通过分解证明:仅匹配边缘分布无法保证风格与类别的独立性。
- 实验显示风格变量可准确预测类别(74%–100%准确率),即便边缘分布接近高斯。
- 提出后处理方案提升生成质量,但跨数据集效果差异大,警示评估局限性。
因子化生成模型常通过将潜风格变量 z_s 的边缘分布匹配到固定高斯先验来正则化,并据此认为风格表示与类别信息无关。本文揭示这一解释错误:仅约束边缘分布无法限制条件分布,导致潜风格仍高度预测标签。我们推导出精确分解,表明该不匹配是实现因子化采样的四个必要条件之一,消除它虽必要但不足以达成目标。实证中,案例模型及四个基线在全局 MMD 接近零的同时,线性探测器仍能以 74%–100% 准确率恢复类别(随机水平为 10%)。所提模型聚类准确率达 99.15%,但外部评估的条件生成仅成功 16%。此泄露在模型容量、课程学习、先验几何和监督等六种独立扰动下持续存在。四种缓解策略使探测准确率降至 21%–46%,但未显著改变类内依赖。后验条件先验在 MNIST 上使生成评分达 0.97(无需重训练),但在 CIFAR-10 上仅达 0.41;经验风格库在 CIFAR-10 上达到 0.88。结果表明,仅基于风格潜变量边缘分布的散度无法验证其与类别的独立性,仅报告边缘统计量无法证实因子化模型的常见宣称。
原文摘要 · Abstract (English)
Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information. We show that this interpretation is incorrect. Matching only the marginal distribution places no constraint on the class-conditional distributions, allowing the latent style to remain highly predictive of the label despite appearing perfectly Gaussian in aggregate. We derive an exact decomposition showing that this mismatch is one of four conditions required for factorized sampling, and demonstrate that eliminating it is necessary but not sufficient to obtain the intended factorization. Empirically, our case-study model and four representative latent baselines achieve near-zero global MMD while still allowing a linear probe to recover class labels with 74%--100% accuracy (10% chance level). Our model reaches 99.15% clustering accuracy, whereas externally evaluated class-conditional generation succeeds only 16% of the time. This leakage remains under six independent perturbations involving model capacity, curriculum, prior geometry, and supervision across two datasets. Four mitigation strategies reduce probe accuracy to 21%--46%, although they leave within-class dependence largely unchanged. A post-hoc conditional prior improves externally evaluated class generation to 0.97 on MNIST without retraining but reaches only 0.41 on CIFAR-10, while an empirical style bank achieves 0.88 on CIFAR-10. These results demonstrate that no divergence computed solely on the marginal distribution of the style latent can certify independence from class labels, and that reporting marginal statistics alone does not verify the property commonly claimed in factorized generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。