arXiv:2603.13299cs.LGcs.AI2026-03

让文生图模型的内部机制可解释,支持精准干预与分析。

DreamReader: An Interpretability Toolkit for Text-to-Image Models

  • 将扩散模型解释统一为可组合的表示操作,支持激活提取与干预。
  • 通过轻量白盒干预,成功在图像中注入目标概念并实现跨模型激活拼接。
  • 引入新方法如LoReFT和梯度引导,适用于研究模型内部表征与迁移性。

尽管文本到图像(T2I)扩散模型迅速普及,其因果与表征层面的分析仍零散且多限于孤立探针技术。为此,我们提出DreamReader:一个统一框架,将扩散模型可解释性形式化为可组合的表示操作,涵盖激活提取、因果修补、结构化消融与模块及时间步的激活操控。DreamReader提供模型无关抽象层,支持对扩散架构进行系统性分析与干预。除整合现有方法外,还引入三种新干预原语:(1) 表示微调(LoReFT),用于子空间约束下的内部自适应;(2) 基于MLP探针训练的分类器引导梯度操控;(3) 模块级跨模型映射,用于系统研究跨模态表征的可迁移性。这些机制使我们能借鉴大语言模型的可解释性技术,对T2I模型实施轻量白盒干预。通过受控实验,我们实现了两模型间的激活拼接,并利用LoReFT操纵多个激活单元,可靠地将目标概念注入生成图像。实验以声明式方式定义,通过批量管道执行,支持可复现的大规模分析。多个案例研究表明,源自语言模型的可解释性技术在扩散模型中也能实现有效可控的干预。DreamReader已开源,推动文生图模型可解释性研究。

原文摘要 · Abstract (English)

Despite the rapid adoption of text-to-image (T2I) diffusion models, causal and representation-level analysis remains fragmented and largely limited to isolated probing techniques. To address this gap, we introduce DreamReader: a unified framework that formalizes diffusion interpretability as composable representation operators spanning activation extraction, causal patching, structured ablations, and activation steering across modules and timesteps. DreamReader provides a model-agnostic abstraction layer enabling systematic analysis and intervention across diffusion architectures. Beyond consolidating existing methods, DreamReader introduces three novel intervention primitives for diffusion models: (1) representation fine-tuning (LoReFT) for subspace-constrained internal adaptation; (2) classifier-guided gradient steering using MLP probes trained on activations; and (3) component-level cross-model mapping for systematic study of transferability of representations across modalities. These mechanisms allows us to do lightweight white-box interventions on T2I models by drawing inspiration from interpretability techniques on LLMs. We demonstrate DreamReader through controlled experiments that (i) perform activation stitching between two models, and (ii) apply LoReFT to steer multiple activation units, reliably injecting a target concept into the generated images. Experiments are specified declaratively and executed in controlled batched pipelines to enable reproducible large-scale analysis. Across multiple case studies, we show that techniques adapted from language model interpretability yield promising and controllable interventions in diffusion models. DreamReader is released as an open source toolkit for advancing research on T2I interpretability.

可解释性文生图扩散模型干预分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。