arXiv:2410.22559cs.LGcs.AI2024-10

揭示生成模型中表征解耦的本质与可识别性机制。

Disentanglement as Identifiable Pushforward Factorisation

  • 通过雅可比矩阵奇异值分解,定义解耦为数据密度的可分离因子结构。
  • 证明解耦需满足两个条件(C1-C2),且各因子可被唯一确定(仅排列与符号差异)。
  • 解释β-VAE中后验对角化如何促进解耦,适用于图像生成与表征学习研究者。

我们对平滑生成推前模型(如变分自编码器和生成对抗网络)中的解耦现象进行了形式化刻画。对于生成器/解码器 $g:Z\to X$ 和因子化先验 $p(z)=\prod_i p_i(z_i)$,我们将解耦定义为推前密度 $p_μ= g_\#p$ 按一维“缝合”因子分解,其中每个潜在维度独立控制数据的一个生成因子。我们证明 $p_μ$ 的分解对应于 $g$ 的雅可比矩阵的奇异值分解;解耦等价于对 $g$ 的两个条件(C1-C2);在满足这些条件下,缝合因子可被识别,仅存在排列与符号不确定性。针对高斯(β-)变分自编码器,我们通过一个恒等式表明:后验对角化在期望下促进条件 C1-C2,从而解释了为何解耦随 $β$ 调节而出现。实验在高斯数据、dSprites 和 CelebA 数据集上验证了该机制。

原文摘要 · Abstract (English)

We characterise disentanglement in smooth generative pushforward models, such as in VAEs and GANs. For a generator/decoder $g:Z\to X$ and factorised prior $p(z)=\prod_i p_i(z_i)$, we define disentanglement as factorisation of the pushforward density $p_μ= g_\#p$ into one-dimensional "seam" factors, where each latent dimension controls an independent generative factor of the data. We prove that $p_μ$ factorises according to the SVD of $g$'s Jacobian; that disentanglement equates to two conditions on $g$ (C1-C2); and that under those conditions the seam factors are identifiable, up to permutation and sign. In the particular case of Gaussian ($β$-)VAEs, we show via an identity how diagonal posteriors promote C1-C2, in expectation, explaining why disentanglement arises modulated by $β$. Experiments illustrate this mechanism on Gaussian data, dSprites, and CelebA.

表征学习生成模型解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。