实时交互式操控视频扩散模型内部结构,实现创作与理解的双重突破
Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers

- 将注意力机制与前馈网络变为可实时调节的交互界面
- 支持对扩散步骤、层、提示词、神经元的逐级精细控制
- 适合艺术家探索模型内核美学,也用于解释性研究
生成式视频模型虽已达到极高的视觉保真度,但其仅通过提示词交互的方式限制了创作者的参与感,并遮蔽了模型内部运作过程。我们提出 Real-Time AttentionBender,一个嵌入 DayDream Scope 生态系统的插件工具,可实时操控视频扩散变换器(DiT)全深度结构。该工具将自注意力、交叉注意力及前馈网络暴露为可独立调节的交互表面,支持精确到特定扩散步、DiT 层、提示词标记和隐藏神经元的定位操作。实时响应带来‘材料亲密感’——直观感受特定层与神经元如何塑造生成视频。本工具兼具 XAIxArts 探针与创造性表达工具双重角色,助力发现模型默认表征空间之外的新美学形态。
原文摘要 · Abstract (English)
Generative video models have achieved remarkable visual fidelity, yet their prompt-only interface offers thin creative agency and obscures the model's material process from the artists working with it. We present Real-Time AttentionBender, a tool that extends the practice of network bending across the full depth of the video diffusion transformer (DiT) and brings it into live, interactive generation. Built as a plugin within the DayDream Scope ecosystem and wrapping open-source real-time Wan pipelines, the tool exposes self-attention, cross-attention, and the feed-forward network as independently manipulable surfaces, with targeting down to individual diffusion steps, DiT layers, prompt tokens, and hidden neurons. The immediacy of live manipulation affords what we call "material intimacy" with the model: a responsive, near-mechanistic feel for how specific layers and neurons shape generated video. We position the tool as simultaneously an XAIxArts probe into transformer internals and an expressive instrument for discovering aesthetics outside the model's default representational space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。