arXiv:2505.24360cs.LG2025-05被引 5

用字典学习解析大模型文本到图像生成机制,实现可控图像生成。

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

  • 采用稀疏自编码器与推理时激活分解方法,提取可解释特征。
  • SAE重建残差流嵌入准确,解释力优于MLP神经元。
  • 可通过激活叠加控制图像生成,适合模型可解释性研究者。

稀疏自编码器(SAEs)是一种有前景的方法,用于分解语言模型激活以实现解释与控制,已在视觉变换器图像编码器和小规模扩散模型中取得成功。最近提出的推理时激活分解(ITDA)是字典学习的一种变体,其字典由激活分布中的数据点构成,并通过梯度追踪进行重构。本文将SAEs与ITDA应用于大型文本到图像扩散模型Flux 1,提出一种可视化自动解释管道。结果表明,SAEs能准确重建残差流嵌入,在解释性上优于MLP神经元;利用SAE特征可通过激活加法实现图像生成的定向调控。ITDA的解释性与SAEs相当。

原文摘要 · Abstract (English)

Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transformer image encoders and to small-scale diffusion models. Inference-Time Decomposition of Activations (ITDA) is a recently proposed variant of dictionary learning that takes the dictionary to be a set of data points from the activation distribution and reconstructs them with gradient pursuit. We apply Sparse Autoencoders (SAEs) and ITDA to a large text-to-image diffusion model, Flux 1, and consider the interpretability of embeddings of both by introducing a visual automated interpretation pipeline. We find that SAEs accurately reconstruct residual stream embeddings and beat MLP neurons on interpretability. We are able to use SAE features to steer image generation through activation addition. We find that ITDA has comparable interpretability to SAEs.

可解释性扩散模型字典学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。