可视化扩散模型生成过程中的注意力动态,助力人机协作理解图像生成逻辑。
Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration
- 基于步骤索引的注意力图谱,追踪每一步的词元间注意力变化。
- 在60个结构化提示的基准上发现可复现的注意力演化模式。
- 交互式时空联动视图,适合研究生成机制与优化人机协作流程。
基于扩散的文生图模型能够生成复杂且高度结构化的视觉内容,但语义结构的涌现与演变仍难以解释。现有工作多依赖聚合注意力或标量摘要,将时间变化与图像空间证据分离。为此,我们提出一种用于探索扩散模型中注意力动态的可视化分析框架:通过步骤索引追踪词元级交叉注意力图谱的演进、其时间集中度及空间关系。该方法结合定量指标与数据驱动的阶段识别,在交互式工作流中实现对生成过程中注意力行为的结构化分析。在包含60个结构化提示的Stable-Diffusion类基准上的案例研究揭示了可复现的、可解释的注意力模式,并表明关联的时间-空间视图有助于观察和讨论生成过程,从而支持更高效的人机协作。
原文摘要 · Abstract (English)
Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structure remain difficult to interpret. Many existing workflows rely on aggregated attention or scalar summaries that separate temporal change from image-space evidence. To address this gap, we present a visual analytics framework for exploring attention dynamics in diffusion models: the step-indexed evolution of token-level cross-attention maps, their temporal concentration, and their spatial relationships. Our approach enables structured analysis of attention behavior across generation steps by integrating quantitative measures with data-driven stage identification in an interactive workflow. Case studies on a structured 60-prompt Stable-Diffusion-class benchmark illustrate recurring, interpretable patterns within this setting and show how linked temporal and spatial views facilitate the observation and discussion of generative processes, supporting more effective human-AI collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。