让扩散Transformer生成过程更可解释,揭示其隐含的层次语义。
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
- 用时序感知的稀疏自编码器提取扩散Transformer各步骤的关键特征
- 发现模型在预训练中自然学到3D结构、物体类别等分层语义
- 提升生成可控性与可解释性,适合安全编辑和风格迁移场景
与基于U-Net的扩散架构相比,扩散Transformer(DiTs)是一类强大但尚未充分探索的生成模型。我们提出TIDE——面向可解释扩散Transformer的时序感知稀疏自编码器框架,旨在从DiTs的多个时间步中提取稀疏且可解释的激活特征。TIDE能有效捕捉随时间变化的表征,并揭示DiTs在大规模预训练过程中自然学习到分层语义,如3D结构、物体类别和细粒度概念。实验表明,TIDE在保持合理生成质量的同时,显著增强了模型的可解释性与可控性,支持安全图像编辑和风格迁移等应用。
原文摘要 · Abstract (English)
Diffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE-Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs-a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。