arXiv:2504.00220cs.LGcs.AI2025-04NeurIPS被引 3

从理论出发,揭示扩散模型如何学出可分离的表征。

Can Diffusion Models Disentangle? A Theoretical Perspective

  • 构建理论框架,明确可分离表征的识别条件。
  • 推导出学习分离子空间的样本复杂度边界。
  • 实验证明理论指导的训练策略能提升分离效果。

本文提出一种新的理论框架,用于理解扩散模型如何学习可分离的表征。在该框架下,我们建立了通用可分离潜在变量模型的可识别性条件,分析了训练动态,并推导出可分离潜在子空间模型的样本复杂度上界。为验证理论,我们在多种任务和模态上进行分离性实验,包括潜在子空间高斯混合模型中的子空间恢复、图像着色、图像去噪以及语音转换用于语音分类。实验还表明,基于理论设计的训练策略(如风格引导正则化)能持续提升分离性能。

原文摘要 · Abstract (English)

This paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations. Within this framework, we establish identifiability conditions for general disentangled latent variable models, analyze training dynamics, and derive sample complexity bounds for disentangled latent subspace models. To validate our theory, we conduct disentanglement experiments across diverse tasks and modalities, including subspace recovery in latent subspace Gaussian mixture models, image colorization, image denoising, and voice conversion for speech classification. Additionally, our experiments show that training strategies inspired by our theory, such as style guidance regularization, consistently enhance disentanglement performance.

扩散模型表征学习可分离性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。