将扩散模型与变分自编码器结合,让无监督表示学习更可解释。
Disentangled representations via score-based variational autoencoders
- 用得分引导的扩散过程优化变分自编码器目标函数
- 在合成数据中恢复真实生成因子,自然图像显式分解语义维度
- 无需标签即可识别有意义的潜在轴,适合探索性分析
我们提出基于得分的多尺度推断自编码器(SAMI),一种将扩散模型与变分自编码器(VAE)理论框架统一的无监督表示学习方法。通过整合两者的证据下界,SAMI 构建了一个严谨的目标函数,利用扩散过程的得分指导来学习数据表示。结果表明,该方法能自动捕捉数据中的有意义结构:在合成数据中恢复真实生成因子,在复杂自然图像中学习到因子化且具有语义意义的潜在维度,并将视频序列编码为比其他编码器更平直的潜在轨迹,尽管仅在静态图像上训练。此外,SAMI 可以在几乎不额外训练的情况下从预训练扩散模型中提取有用表示。其显式的概率形式提供了在无监督条件下识别语义轴的新方法,数学上的精确性使我们能够对所学表示的本质做出形式化陈述。总体而言,这些结果表明,扩散模型中的隐含结构性信息可通过与变分自编码器协同,转化为显式且可解释的表示。
原文摘要 · Abstract (English)
We present the Score-based Autoencoder for Multiscale Inference (SAMI), a method for unsupervised representation learning that combines the theoretical frameworks of diffusion models and VAEs. By unifying their respective evidence lower bounds, SAMI formulates a principled objective that learns representations through score-based guidance of the underlying diffusion process. The resulting representations automatically capture meaningful structure in the data: it recovers ground truth generative factors in our synthetic dataset, learns factorized, semantic latent dimensions from complex natural images, and encodes video sequences into latent trajectories that are straighter than those of alternative encoders, despite training exclusively on static images. Furthermore, SAMI can extract useful representations from pre-trained diffusion models with minimal additional training. Finally, the explicitly probabilistic formulation provides new ways to identify semantically meaningful axes in the absence of supervised labels, and its mathematical exactness allows us to make formal statements about the nature of the learned representation. Overall, these results indicate that implicit structural information in diffusion models can be made explicit and interpretable through synergistic combination with a variational autoencoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。