用层叠图理论解析大模型语义超位置,揭示解释性失效的三类根源。
A Gauge Theory of Superposition: Toward a Sheaf-Theoretic Atlas of Neural Representations
- 将局部语义空间与信息几何结合,构建可计算的非局域干扰度量。
- 实证发现:干扰能量有界且零误报,剪枝后仍稳定收敛。
- 适合研究模型可解释性、神经表征结构的学者深度阅读。
我们为大型语言模型中的超位置现象构建了一个离散规范场理论框架,取代单一全局词典假设,转而采用局部语义图谱构成的层叠图谱。上下文被聚类为分层上下文复形,每个图谱携带局部特征空间和局部信息几何度量(Fisher/高斯-牛顿),用于识别预测关键的特征交互。该框架导出一个Fisher加权的干扰能量,并揭示三类全局可解释性的可观测障碍:(O1) 局部拥堵(活跃负载超过Fisher带宽)、(O2) 代理剪切(几何传输与固定对应代理不匹配)、(O3) 非平凡全同性(路径依赖的环路传输)。在冻结的Llama-3.2-3B Instruct模型上,基于WikiText-103、C4衍生英文网络文本子集及the-stack-smol数据集,验证了四项结果:(A) 在生成树上进行构造性规范固定后,每条弦残差等于其基本圈的全同性,使全同性可计算且规范不变;(B) 剪切下界给出数据依赖的传输错配能量,使$D_{\mathrm{shear}}$成为不可避免的失败界限;(C) 获得高覆盖率、零违规的非平凡认证拥堵/干扰边界;(D) 自举与样本量实验显示$D_{\mathrm{shear}}$与$D_{\mathrm{hol}}$估计稳定,良好条件子系统中收敛性更优。
原文摘要 · Abstract (English)
We develop a discrete gauge-theoretic framework for superposition in large language models (LLMs) that replaces the single-global-dictionary premise with a sheaf-theoretic atlas of local semantic charts. Contexts are clustered into a stratified context complex; each chart carries a local feature space and a local information-geometric metric (Fisher/Gauss-Newton) identifying predictively consequential feature interactions. This yields a Fisher-weighted interference energy and three measurable obstructions to global interpretability: (O1) local jamming (active load exceeds Fisher bandwidth), (O2) proxy shearing (mismatch between geometric transport and a fixed correspondence proxy), and (O3) nontrivial holonomy (path-dependent transport around loops). We prove and instantiate four results on a frozen open LLM (Llama-3.2-3B Instruct) using WikiText-103, a C4-derived English web-text subset, and the-stack-smol. (A) After constructive gauge fixing on a spanning tree, each chord residual equals the holonomy of its fundamental cycle, making holonomy computable and gauge-invariant. (B) Shearing lower-bounds a data-dependent transfer mismatch energy, turning $D_{\mathrm{shear}}$ into an unavoidable failure bound. (C) We obtain non-vacuous certified jamming/interference bounds with high coverage and zero violations across seeds and hyperparameters. (D) Bootstrap and sample-size experiments show stable estimation of $D_{\mathrm{shear}}$ and $D_{\mathrm{hol}}$, with improved concentration on well-conditioned subsystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。