无需训练即可精准定位图像中的文本概念,提升视觉任务可解释性。
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
- 利用扩散变压器注意力层参数生成上下文感知的概念嵌入
- 输出空间线性投影使显著图清晰度显著优于传统交叉注意力
- 零样本图像分割性能领先15种方法,适用于图像与视频生成
多模态扩散变压器(DiTs)的丰富表示是否具备增强可解释性的独特属性?我们提出ConceptAttention,一种新方法,通过复用DiT注意力层参数,在不需额外训练的前提下生成高质量显著图,精准定位图像中的文本概念。核心发现是:在DiT注意力层输出空间进行线性投影,能产生远比常见交叉注意力图更清晰的显著图。该方法在零样本图像分割基准上表现优异,超越15种现有零样本可解释性方法,在ImageNet-Segmentation数据集上达到最先进水平。ConceptAttention适用于主流图像模型,并可无缝扩展至视频生成任务。本工作首次提供证据表明,多模态DiTs的表示具有高度迁移性,可有效用于分割等视觉任务。
原文摘要 · Abstract (English)
Do the rich representations of multi-modal diffusion transformers (DiTs) exhibit unique properties that enhance their interpretability? We introduce ConceptAttention, a novel method that leverages the expressive power of DiT attention layers to generate high-quality saliency maps that precisely locate textual concepts within images. Without requiring additional training, ConceptAttention repurposes the parameters of DiT attention layers to produce highly contextualized concept embeddings, contributing the major discovery that performing linear projections in the output space of DiT attention layers yields significantly sharper saliency maps compared to commonly used cross-attention maps. ConceptAttention even achieves state-of-the-art performance on zero-shot image segmentation benchmarks, outperforming 15 other zero-shot interpretability methods on the ImageNet-Segmentation dataset. ConceptAttention works for popular image models and even seamlessly generalizes to video generation. Our work contributes the first evidence that the representations of multi-modal DiTs are highly transferable to vision tasks like segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。