揭示音视频模型中跨模态注意力的泄漏机制
The Attention Triangle in Audio-Video Models

- 分析三模态注意力三角,发现音视频双向影响
- 当提示与模型先验冲突时,易产生错误视觉输出
- 提出诊断信号,可控制性诱发或抑制语义泄漏
音视频扩散模型依赖跨模态注意力协调文本、声音和视觉内容,但该机制可能引入细微且系统性的语义泄漏。我们通过探测和分析由文本-音频、音频-视频、文本-视频三组跨模态注意力边构成的“注意力三角”,研究生成过程中语义信息在模态间的传递路径。分析发现,音频-视频边呈现双向影响:音频可影响视频生成,视频也可影响音频生成。该路径受模型参数中编码的偏差塑造,成为泄漏的主要来源:当输入提示与模型学习到的先验存在矛盾时,跨模态交互可能覆盖预期条件,将语义重定向至视觉上典型但错误的结果。这些现象表明,语义伪影并非仅源于注意力扩散,而是特定路径上的结构化、偏见驱动的交互所致。基于此视角,我们提取注意力衍生信号,用于暴露语义在各模态中的分布与锚定状态,并作为诊断工具,在受控条件下分析或刻意引发泄漏。这使我们能够探测跨模态路由的内部动态,并分离单个交互的作用。进一步利用这些信号,设计推理时干预策略,促进更一致的跨模态对齐。大量实验验证了分析的有效性,证明在保持生成质量的同时提升了语义锚定准确性。
原文摘要 · Abstract (English)
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。