揭示注意力谱诊断的局限性,指出其无法捕捉信息流向方向。
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
- 将注意力算子分解为对称与反对称部分,分别对应传输能力和方向性
- 证明所有转置不变的谱诊断均忽略信息流向方向,敏感度受不对称系数限制
- 提出双轴诊断框架,可预测不同模型在幻觉检测中的方向性表现
每个注意力头定义了一个度归一化的传输算子,越来越多的诊断方法通过其谱来解读模型行为(如幻觉)。我们探讨此类诊断能与不能推断的内容。该算子可正交分解为决定传输能力的对称部分和编码方向性的反对称部分。我们证明了可辨识性极限:所有转置不变的谱诊断均为方向盲区(无法区分算子与其转置,故忽略信息流方向),且任意Lipschitz诊断的转置敏感度被不对称系数 $G$ 所限制。这约束了谱诊断对注意力算子的分辨能力(如 LapEigvals 和 LLM-Check 的注意力谱分支)。在剩余轴上,闭式双分图-Cheeger景观显示:均匀因果注意力遵循与 $n$ 无关的 $ϕ\ge 1/5$ 时间切割下界,而窗口注意力可降至 $O(w/n)$;该下界为理想化基准而非经验吸引子,实际头部低于该值的比例本身是稳定的架构描述符。双轴诊断($ϕ$ 表示容量,$G$ 表示不对称强度)产生可证伪的极性预测,在长度控制、强制评分评估下于解码器仅用、编码器仅用及编码器-解码器模型中得到验证(容量轴信号 0.62–0.84 LC-AUROC):HalucEval 与 MedHallu 中极性反转,方向符合预测但强度不对称,决策极性按模型范式校准。
原文摘要 · Abstract (English)
Every attention head defines a degree-normalized transport operator, and a growing family of diagnostics reads model behavior (hallucination among them) from its spectrum. We ask what such diagnostics can and cannot infer. The operator splits orthogonally into a symmetric part governing transport \emph{capacity} and an antisymmetric part encoding \emph{orientation}. We prove an identifiability limit: every transpose-invariant spectral diagnostic is \emph{orientation-blind} (unable to distinguish an operator from its transpose, hence blind to the orientation of information flow), with a transpose-stability bound limiting any Lipschitz diagnostic's transpose sensitivity by the asymmetry coefficient $G$. This bounds what spectral diagnostics of the attention operator can resolve (e.g.\ LapEigvals and the attention-spectral branch of LLM-Check). On the surviving axis, a closed-form bipartite-Cheeger landscape shows uniform causal attention obeys an $n$-independent \emph{temporal-cut} floor $ϕ\ge 1/5$ while window attention pierces it as $O(w/n)$; the floor is an idealized benchmark, not an empirical attractor, and the fraction of real heads falling below it is itself an empirically stable architectural descriptor. The two-axis diagnostic ($ϕ$ for capacity, $G$ for asymmetry magnitude) yields a falsifiable polarity prediction, borne out \emph{in sign} under length-controlled, forced-scoring evaluation across decoder-only, encoder-only, and encoder--decoder models (capacity-axis signal 0.62--0.84 LC-AUROC): polarity reverses between HaluEval and MedHallu, directionally as predicted though asymmetric in strength, with decision polarity calibrated per regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。