揭示Stable Diffusion latent空间中颜色的编码规律。
Color encoding in Latent Space of Stable Diffusion Models
- 通过可控数据集与PCA分析,发现颜色沿圆形对立轴编码。
- 颜色主要由c_3、c_4通道表示,强度与形状在c_1、c_2通道。
- 结果有助于模型理解与可控生成,适合研究者参考。
基于扩散的生成模型虽已实现高视觉保真度,但对颜色、形状等感知属性在内部如何表征仍缺乏深入理解。本文通过系统分析Stable Diffusion的潜在表示,结合可控合成数据集、主成分分析(PCA)与相似性度量,发现颜色信息主要沿圆形对立轴编码,集中在潜空间通道c_3和c_4;而强度与形状则主要由通道c_1和c_2表示。结果表明,Stable Diffusion的潜空间具有符合高效编码理论的可解释结构。这些发现为模型理解、编辑应用及更解耦生成框架的设计奠定了基础。
原文摘要 · Abstract (English)
Recent advances in diffusion-based generative models have achieved remarkable visual fidelity, yet a detailed understanding of how specific perceptual attributes - such as color and shape - are internally represented remains limited. This work explores how color is encoded in a generative model through a systematic analysis of the latent representations in Stable Diffusion. Through controlled synthetic datasets, principal component analysis (PCA) and similarity metrics, we reveal that color information is encoded along circular, opponent axes predominantly captured in latent channels c_3 and c_4, whereas intensity and shape are primarily represented in channels c_1 and c_2. Our findings indicate that the latent space of Stable Diffusion exhibits an interpretable structure aligned with a efficient coding representation. These insights provide a foundation for future work in model understanding, editing applications, and the design of more disentangled generative frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。