用残差建模扩散模型激活轨迹,揭示随时间演化的语义特征。
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

- 基于相邻时刻的线性预测,用残差表示激活轨迹
- 稀疏编码器学习到的潜变量可对应时序特征路径
- 适合研究扩散模型内部动态,支持可控图像生成
文本到图像扩散模型通过迭代去噪过程生成图像,其内部神经层产生的是随时间演化的激活轨迹而非静态表示。近期研究使用稀疏自编码器(SAEs)将扩散激活分解为可解释的特征方向,但多数方法仅分析单个时间步或依赖时间条件,而非直接从完整激活轨迹中学习。本文提出残差化时序稀疏自编码器(residualized temporal SAEs),收集去噪过程中各时间步的激活,拟合相邻时间步间的线性预测关系,以初始激活加未被线性动态解释的残差分量表示每条轨迹。在该残差表示上训练SAE,促使稀疏潜变量捕捉线性可预测之外的结构。残差解码方向可映射回激活空间,使每个潜变量可被分析为去噪过程中的特征演化轨迹。通过重建、消融实验、时空特征分析及对Stable Diffusion 1.5的定性控制实验,验证了该方法能有效研究具有时序结构的扩散激活。
原文摘要 · Abstract (English)
Text-to-image diffusion models generate images through an iterative denoising process, so internal neural layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable feature directions, but most approaches analyze activations at individual timesteps or condition on time rather than learning directly from full activation trajectories. In this work, we introduce residualized temporal SAEs for diffusion activation trajectories. We collect activations across denoising time, fit linear predictors between neighboring timesteps, and represent each trajectory using an initial activation together with residual components not explained by these linear dynamics. Training an SAE on this residualized representation encourages sparse latents to capture structure beyond what is linearly predictable. The residualized decoder directions can be mapped back into activation space, allowing each latent to be analyzed as a feature trajectory over denoising time. Through reconstruction and ablation studies, spatiotemporal feature analysis, and qualitative steering experiments on Stable Diffusion~1.5, we show that residualized temporal SAEs provide a useful framework for studying temporally structured diffusion activations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。