用扩散模型提升表征解耦效果,让每个隐变量只控制一个属性。
Disentangling Disentangled Representations: Towards Improved Latent Units via Diffusion Models
- 设计动态高斯锚定机制,明确划分不同属性的隐变量边界。
- 在合成与真实数据上达到当前最优解耦性能,下游任务表现更优。
- 提出跳过丢弃技术,让扩散模型更适配解耦特征提取器。
解耦表征学习(DRL)旨在将观测数据分解为内在核心因素,以深入理解数据。在现实场景中,手动定义和标注这些因素非常困难,因此无监督方法更具吸引力。近期对扩散模型(DMs)在无监督DRL中的应用探索有限,而扩散模型本身具有特定归纳偏置,可确保输入扩散模型的每个隐变量仅表达单一特征。为此,本文设计了动态高斯锚定机制,强制使隐变量在属性上分离,增强可解释性,并促进隐变量间的独立性。此外,提出跳过丢弃技术,通过简单修改去噪U-Net结构,使其更适应解耦特征提取器,克服其与解耦结构不兼容的问题。所提方法兼顾隐变量语义与扩散模型结构特性,显著提升基于扩散模型的解耦表征实用性,在合成与真实数据上均取得当前最优解耦性能,并在下游任务中展现优势。
原文摘要 · Abstract (English)
Disentangled representation learning (DRL) aims to break down observed data into core intrinsic factors for a profound understanding of the data. In real-world scenarios, manually defining and labeling these factors are non-trivial, making unsupervised methods attractive. Recently, there have been limited explorations of utilizing diffusion models (DMs), which are already mainstream in generative modeling, for unsupervised DRL. They implement their own inductive bias to ensure that each latent unit input to the DM expresses only one distinct factor. In this context, we design Dynamic Gaussian Anchoring to enforce attribute-separated latent units for more interpretable DRL. This unconventional inductive bias explicitly delineates the decision boundaries between attributes while also promoting the independence among latent units. Additionally, we also propose Skip Dropout technique, which easily modifies the denoising U-Net to be more DRL-friendly, addressing its uncooperative nature with the disentangling feature extractor. Our methods, which carefully consider the latent unit semantics and the distinct DM structure, enhance the practicality of DM-based disentangled representations, demonstrating state-of-the-art disentanglement performance on both synthetic and real data, as well as advantages in downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。