arXiv:2509.05323cs.AIcs.MM2025-09被引 1

用注意力图解构AI视频生成,让艺术创作看得见。

Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts

  • 提取并可视化文本到视频生成中的交叉注意力图
  • 揭示了模型在时空上的注意力分布模式
  • 适合对AI艺术创作与可解释性感兴趣的创作者

本文对视频扩散变换器的注意力机制进行了艺术与技术双重探索。受早期视频艺术家通过操控模拟信号创造新视觉美学的启发,本研究提出一种从生成视频模型中提取和可视化交叉注意力图的方法。基于开源Wan模型,该工具为文本到视频生成过程中的时间与空间注意力行为提供了可解释窗口。通过探索性分析与艺术案例研究,我们考察了注意力图作为分析工具及原始艺术素材的潜力。这项工作推动了面向艺术的可解释人工智能(XAIxArts)领域的发展,邀请艺术家重新掌握AI内部运作作为创作媒介。

原文摘要 · Abstract (English)

This paper presents an artistic and technical investigation into the attention mechanisms of video diffusion transformers. Inspired by early video artists who manipulated analog video signals to create new visual aesthetics, this study proposes a method for extracting and visualizing cross-attention maps in generative video models. Built on the open-source Wan model, our tool provides an interpretable window into the temporal and spatial behavior of attention in text-to-video generation. Through exploratory probes and an artistic case study, we examine the potential of attention maps as both analytical tools and raw artistic material. This work contributes to the growing field of Explainable AI for the Arts (XAIxArts), inviting artists to reclaim the inner workings of AI as a creative medium.

视频生成注意力机制AI艺术可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。