提出新型扩散自编码器,兼顾高效生成与强表征能力
On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning
- 设计可学习噪声过程的隐变量,实现输入相关建模
- 在下游任务中表现优于标准扩散模型,生成步数更少
- 适合需要高效生成和鲁棒表征的场景
扩散自编码器(DAs)通过输入依赖的隐变量,在扩散过程中同时捕捉表征,可用于分类、可控生成和插值等任务。然而其生成性能高度依赖隐变量建模质量。另一类扩散模型通过学习前向(加噪)过程来提升建模效果,但需满足扩散过程终态约束。本文揭示两类模型的联系,提出一种新设计(称作DMZ):通过优化隐变量选择与条件机制,兼顾下游任务表现(如领域迁移)与生成效率,显著减少去噪步骤,优于标准扩散模型。
原文摘要 · Abstract (English)
Diffusion autoencoders (DAs) are variants of diffusion generative models that use an input-dependent latent variable to capture representations alongside the diffusion process. These representations, to varying extents, can be used for tasks such as downstream classification, controllable generation, and interpolation. However, the generative performance of DAs relies heavily on how well the latent variables can be modelled and subsequently sampled from. Better generative modelling is also the primary goal of another class of diffusion models -- those that learn their forward (noising) process. While effective at adjusting the noise process in an input-dependent manner, they must satisfy additional constraints derived from the terminal conditions of the diffusion process. Here, we draw a connection between these two classes of models and show that certain design decisions (latent variable choice, conditioning method, etc.) in the DA framework -- leading to a model we term DMZ -- allow us to obtain the best of both worlds: effective representations as evaluated on downstream tasks, including domain transfer, as well as more efficient modelling and generation with fewer denoising steps compared to standard DMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。