arXiv:2605.25225cs.LGcs.AI2026-05

用场论视角解析Transformer,让干预实验有数学依据。

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

论文配图:Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability
图 1 · 摘自论文原文
  • 把残差流看作层深与位置的场,干预变作场中局部源注入。
  • 一阶敏感度可预测不同位置的干预效果,传播路径呈各向异性。
  • 适合研究模型内部机制的人,尤其是想系统化理解干预实验者。

机械可解释性常通过激活修补、因果追踪等方法干预Transformer内部激活。本文提出Transformer场论:将固定前向传播中的残差流视为层深度与标记位置上的场。在此框架下,修补操作等价于在场中引入局部源,一阶敏感度场可预测修补效应,格林函数描述下游传播,而修补选择被建模为伴随逆问题。实证上,我们在GPT-2风格的自回归Transformer中测试了该理论的前向响应对象。结果显示:局部场干预存在有限线性区域;一阶敏感度能准确预测跨层-标记位点的修补效果;局部源引发结构化的各向异性场传播;高敏感度区域与截断格林算子可提供简化响应描述;提示引发的场位移部分传递答案行为。这些结果确立了敏感度、场响应与截断格林算子作为组织修补实验的实用对象,并为补丁位点推断与跨尺度响应迁移提供了正向数学基础。

原文摘要 · Abstract (English)

Mechanistic interpretability often studies Transformer behavior by intervening on internal activations through activation patching, causal tracing, path patching, and steering directions. This paper develops Transformer Field Theory: a response-theoretic framework in which the residual stream of a fixed forward pass is treated as a Transformer field over layer depth and token position. In this formulation, patching becomes a localized source insertion into the Transformer field, first-order sensitivity fields predict patch effects, Green functions describe downstream propagation, and patch selection is posed as an adjoint inverse problem. Empirically, we test the theory's forward response objects in GPT-2-style autoregressive Transformers. Localized Transformer-field interventions exhibit a bounded local linear regime; first-order sensitivities predict patch effects across layer-token sites; localized sources generate structured anisotropic Transformer-field propagation; high-sensitivity sites and sliced Green operators provide reduced response descriptions; and prompt-induced Transformer-field displacements partially transfer answer behavior. These results establish sensitivities, Transformer-field responses, and sliced Green operators as practical objects for organizing patching experiments, while providing the forward mathematical basis for patch-site inference and cross-scale response transfer.

可解释性Transformer场论干预分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。