arXiv:2604.20936cs.MMcs.CV2026-04被引 1

通过操控注意力图,让艺术家探索视频生成模型的内部机制。

AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe

论文配图:AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
图 1 · 摘自论文原文
  • 用2D变换修改跨注意力图,实现对生成过程的干预
  • 4500+次实验显示注意力高度耦合,局部控制难实现
  • 既可解释模型机制,也能生成独特视觉风格

我们提出 AttentionBender,一个用于操纵视频扩散变换器中跨注意力机制的工具,帮助艺术家探究黑箱视频生成模型的内部运作。尽管生成结果日益逼真,但仅靠提示词控制限制了艺术家对模型生成过程的理解,也难以突破其默认倾向。基于一种自传式的设计研究方法,我们在 Network Bending 基础上构建 AttentionBender,通过对跨注意力图施加二维变换(如旋转、缩放、平移等)来调节生成效果。我们通过在不同提示词、操作类型和层目标下可视化超过4500个视频生成结果进行评估。结果显示,跨注意力高度耦合:针对性的修改往往无法实现干净的局部控制,反而引发分布式的失真与故障美学,而非线性编辑。AttentionBender 不仅作为可解释AI的注意力机制探针,还可作为一种创造新视觉风格的技术,拓展模型学习表征空间之外的表现力。

原文摘要 · Abstract (English)

We present AttentionBender, a tool that manipulates cross-attention in Video Diffusion Transformers to help artists probe the internal mechanics of black-box video generation. While generative outputs are increasingly realistic, prompt-only control limits artists' ability to build intuition for the model's material process or to work beyond its default tendencies. Using an autobiographical research-through-design approach, we built on Network Bending to design AttentionBender, which applies 2D transforms (rotation, scaling, translation, etc.) to cross-attention maps to modulate generation. We assess AttentionBender by visualizing 4,500+ video generations across prompts, operations, and layer targets. Our results suggest that cross-attention is highly entangled: targeted manipulations often resist clean, localized control, producing distributed distortions and glitch aesthetics over linear edits. AttentionBender contributes a tool that functions both as an Explainable AI style probe of transformer attention mechanisms, and as a creative technique for producing novel aesthetics beyond the model's learned representational space.

视频生成注意力机制创意探针扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。