arXiv:2606.10046cs.SDcs.AI2026-06

破解音频分离模型的注意力机制,实现加速且不降质。

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

论文配图:Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models
图 1 · 摘自论文原文
  • 通过因果干预分析注意力动态,发现双路径文本控制机制。
  • 稳定层早期构建时间骨架,快速层持续消除噪声伪影。
  • 提出层选择性缓存法,推理加速6.7倍且质量损失极小。

Flow-matching transformers 在音频分离任务中表现优异,但其注意力机制不透明。本文将因果干预原理转化为适用于 SAM Audio 的确定性推理阶段探测方法。正交探测揭示了双路径文本条件机制:加性注入控制语义身份,交叉注意力优化声学结构。观察到分层异步收敛现象:稳定层早期建立时间框架,快速层在采样过程中持续修正伪影。模型还抑制时间分割线索以维持连续流稳定性。基于这些发现,提出无需训练的层选择性注意力缓存(LSAC)方法,在不同声学复杂度下减少约25%的自注意力计算量,质量损失微乎其微,且相比简单步数减少,质量保留提升高达6.7倍。

原文摘要 · Abstract (English)

Flow-matching transformers achieve strong audio separation, yet their attention dynamics are opaque. We adapt established causal-intervention principles into a deterministic, inference-time probing protocol for SAM Audio. Orthogonal probing uncovers a dual-pathway text-conditioning mechanism: additive injections control semantic identity, while cross-attention refines acoustic structure. We observe an asynchronous layerwise convergence: stable layers build temporal scaffolds early, whereas fast layers continue resolving artifacts during sampling. The model also attenuates temporal segmentation cues to maintain continuous-flow stability. Using these insights, we propose Layer-Selective Attention Caching (LSAC), a training-free acceleration method that caches attention in stable layers. Across acoustic complexities, LSAC cuts self-attention computation by about ~25% with negligible quality loss and yields up to 6.7x higher quality retention than naive step reduction.

音频分离注意力机制推理加速因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。