arXiv:2608.01676cs.CLcs.LG2026-08

揭示稀疏注意力如何改变长文本模型对内容的依赖,发现压缩会放大或反转关键信息的影响。

Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation

论文配图:Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation
图 1 · 摘自论文原文
  • 设计对抗性探针实验,分离出稀疏化对内容影响的真实效应
  • 在高压缩下,3/4场景出现关键信息影响力增强甚至反转
  • 适用于评估长文本模型中注意力机制可靠性,尤其关注安全敏感任务

稀疏注意力广泛用于长上下文模型部署,但缺乏框架来评估丢弃块如何改变特定内容对输出的影响。我们首次证明该现象真实且具有因果性:在四种架构上,块稀疏闪注意力(BSFA)路径重播导致16个单元中有13个输出决策发生变化,且零身份重播标签翻转。我们引入一种密集校准的反事实审计方法,使用匹配的探针卡——黄金(正确答案)、毒药(目标错误答案)和良性(仅填充)——在六种布局对称条件下,隔离稀疏化的特异性影响。两种模式竞争:信号集中性表现为选择保留黄金与毒药块远高于匹配填充的良性块(G≈P≫B);整合损失表现为丢弃块会切断跨块注意力,实验证明孤立探针块影响从4.48对数几率降至零。压缩比决定平衡:在四组模型-任务对中,从轻度(c=0.25)到重度(c=0.75)压缩的全范围测试显示,三例在高压缩下趋向更强的稀疏放大,两例发生符号反转。三种独立手段——BSFA路径重播、受控块top-k、KV缓存淘汰——结果一致:稀疏化以聚合准确率无法察觉的方式改变了内容影响力。我们提供一个开放可部署的测量框架,适用于任何暴露块身份的模型。

原文摘要 · Abstract (English)

Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output. We first establish that the phenomenon is real and causal: Block Sparse Flash Attention (BSFA) route replay across four architectures changes output decisions in 13 of 16 cells, with zero identity-replay label flips. We then introduce a dense-calibrated counterfactual audit using matched probe cards---Gold (carrying the correct answer label), Poison (carrying a target wrong label), and Benign (filler only)---under six-layout position symmetry, isolating the sparsification-specific effect. Two patterns compete. Signal concentration: the selector preserves Gold and Poison blocks far above filler-matched Benign blocks (G$\approx$P$\gg$B across all model--task pairs). Integration loss: discarding blocks severs cross-block attention---confirmed by an ablation where isolating the probe block collapses its influence from 4.48 logits to zero. Compression ratio governs the balance: a full sweep from mild ($c=0.25$) to aggressive ($c=0.75$) compression across four model--task pairs reveals that three of four cells move toward stronger sparse amplification at higher compression, with two exhibiting sign reversals. Three independent arms---BSFA route replay, controlled block-top-$k$, and KV-cache eviction---converge: sparsification changes content influence in ways aggregate accuracy cannot detect. We provide an open measurement framework deployable on any model exposing block identities.

稀疏注意力长文本建模模型可解释性反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。