arXiv:2602.10216cs.LG2026-02被引 2

让扩散模型的生成结果可控,精准调控图像细节。

ELROND: Exploring and decomposing intrinsic capabilities of diffusion models

  • 在输入嵌入空间中分解出可解释的语义方向。
  • 提升模型多样性,缓解蒸馏后的模式崩溃问题。
  • 通过子空间维度评估概念复杂度,适合视觉控制研究者。

单一文本提示输入扩散模型后,由于随机过程的影响,常产生多种视觉输出,用户难以控制具体语义变化。现有无监督方法虽分析输出特征,却忽略生成过程本身。本文提出框架,在输入嵌入空间直接解耦这些语义方向:通过反向传播固定提示下不同随机实现间的差异,获取梯度,并利用主成分分析或稀疏自编码器将其分解为有意义的引导方向。该方法带来三项贡献:(1)分离出可解释、可操控的语义方向,实现对单个概念的精细控制;(2)通过重引入丢失的多样性,有效缓解蒸馏模型中的模式崩溃;(3)基于发现子空间的维度,建立一种新型概念复杂度估计器,适用于特定模型。

原文摘要 · Abstract (English)

A single text prompt passed to a diffusion model often yields a wide range of visual outputs determined solely by stochastic process, leaving users with no direct control over which specific semantic variations appear in the image. While existing unsupervised methods attempt to analyze these variations via output features, they omit the underlying generative process. In this work, we propose a framework to disentangle these semantic directions directly within the input embedding space. To that end, we collect a set of gradients obtained by backpropagating the differences between stochastic realizations of a fixed prompt that we later decompose into meaningful steering directions with either Principal Components Analysis or Sparse Autoencoder. Our approach yields three key contributions: (1) it isolates interpretable, steerable directions for precise, fine-grained control over a single concept; (2) it effectively mitigates mode collapse in distilled models by reintroducing lost diversity; and (3) it establishes a novel estimator for concept complexity under a specific model, based on the dimensionality of the discovered subspace.

扩散模型可控生成语义解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。