arXiv:2606.12917cs.LG2026-06中稿 · ICML被引 1

首次揭示表格大模型中注意力头的计算分工与动态变化规律。

Where Computation Lives Inside TabPFN: Causal Localisation of Attention Head Function

论文配图:Where Computation Lives Inside TabPFN: Causal Localisation of Attention Head Function
图 1 · 摘自论文原文
  • 通过激活修补等方法,发现注意力头在不同层有明确分工。
  • 一个核心注意力头在峰值层的贡献是其他头的2至5倍。
  • 适合研究模型机制与可解释性的研究人员阅读。

我们首次对表格基础模型TabPFN 2.5进行了因果机制分析,研究其特征级注意力头在各层的计算分布。基于两种合成回归数据集,利用激活修补、消融实验及注意力熵分析,发现显著的时间特化现象:一个注意力头的因果必要性在峰值层高出其他头2至5倍,且其主导层随任务复杂度变化而转移;其余头则呈现对称的晚期层表现。注意力熵与修补结果相互印证了主导头的活跃层。此外,我们尝试通过对比激活引导实现推理时可调控性,但该方法无法跨样本迁移。我们将其归因于TabPFN的上下文学习机制——任务结构通过上下文依赖的注意力编码,而非语言模型中稳定的参数方向,导致难以实现有效引导。

原文摘要 · Abstract (English)

We present the first causal mechanistic analysis of a tabular foundation model, investigating how TabPFN 2.5's feature wise attention heads distribute computation across layers. Using activation patching, ablation, and attention entropy across two synthetic regression datasets, we find clear temporal specialisation: one head's causal necessity dominates that of the others by 2 to 5 times at peak layer, with its dominant layer shifting across tasks of different complexity, while the remaining heads exhibit symmetric late layer profiles. Attention entropy and patching provide convergent evidence for the computationally active layers of the dominant head. We additionally investigate inference time steerability via contrastive activation steering, which fails to transfer across samples. We attribute this result to TabPFN's in context learning mechanism, which encodes task structure through context dependent attention rather than the stable parametric directions that make steering tractable in language models.

模型解释注意力机制表格模型因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。