arXiv:2505.05732cs.LGcs.CV2025-05

用扩散模型学语义嵌入,效果超越主流自监督方法

Automated Learning of Semantic Embedding Representations for Diffusion Models

  • 设计多级去噪自编码框架,通过自条件扩散学习获取时序一致嵌入
  • 在多个数据集上嵌入质量显著优于现有自监督方法,提升明显
  • 适合需要高质量语义表示的通用深度学习任务

生成模型能捕捉数据的真实分布,生成语义丰富的表示。去噪扩散模型(DDMs)具备优异的生成能力,但其高效的表征学习方法仍不足。本文提出一种多层级去噪自编码框架,引入顺序一致的扩散变换器和时间步依赖的编码器,通过自条件扩散学习,在去噪马尔可夫链上获取嵌入表示。直观上,该编码器在不同噪声水平下将高维数据压缩为潜在空间中的方向向量,实现跨所有时间步的图像嵌入学习。在多个数据集上的大量实验表明,经优化的扩散模型嵌入在大多数情况下超越当前最先进的自监督表征学习方法,展现出卓越的判别性语义表示能力。本工作证明了扩散模型不仅适用于生成任务,也具备应用于通用深度学习任务的潜力。

原文摘要 · Abstract (English)

Generative models capture the true distribution of data, yielding semantically rich representations. Denoising diffusion models (DDMs) exhibit superior generative capabilities, though efficient representation learning for them are lacking. In this work, we employ a multi-level denoising autoencoder framework to expand the representation capacity of DDMs, which introduces sequentially consistent Diffusion Transformers and an additional timestep-dependent encoder to acquire embedding representations on the denoising Markov chain through self-conditional diffusion learning. Intuitively, the encoder, conditioned on the entire diffusion process, compresses high-dimensional data into directional vectors in latent under different noise levels, facilitating the learning of image embeddings across all timesteps. To verify the semantic adequacy of embeddings generated through this approach, extensive experiments are conducted on various datasets, demonstrating that optimally learned embeddings by DDMs surpass state-of-the-art self-supervised representation learning methods in most cases, achieving remarkable discriminative semantic representation quality. Our work justifies that DDMs are not only suitable for generative tasks, but also potentially advantageous for general-purpose deep learning applications.

扩散模型语义嵌入自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。