arXiv:2504.18274cs.LG2025-04被引 11

用物理方法解析小模型,找出关键注意力头的功能归属。

Structural Inference: Interpreting Small Language Models with Susceptibilities

  • 将神经网络视作统计力学系统,通过数据扰动探测响应
  • 300万参数模型中识别出多词模式与归纳推理模块
  • 适合研究模型可解释性或注意力机制的开发者

我们提出一种线性响应框架用于可解释性分析,将神经网络视为贝叶斯统计力学系统。对数据分布施加微小扰动(如将The Pile数据集偏向GitHub代码或法律文本),会引发网络特定组件上可观测量后验期望的一阶变化。该响应称为敏感度,可通过局部SGLD采样高效估计,并分解为带符号的逐标记贡献,作为归因分数。我们将这些敏感度整合为响应矩阵,其低秩结构能有效分离300万参数Transformer中的功能模块,如多词模式头与归纳推理头。

原文摘要 · Abstract (English)

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of an observable localized on a chosen component of the network. The resulting susceptibility can be estimated efficiently with local SGLD samples and factorizes into signed, per-token contributions that serve as attribution scores. We combine these susceptibilities into a response matrix whose low-rank structure separates functional modules such as multigram and induction heads in a 3M-parameter transformer.

可解释性注意力机制小模型统计力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。