arXiv:2605.00333cs.LGcs.CL2026-05

冻结的Gemma模型在非文本任务中表现出独特注意力头重要性,揭示跨模态迁移的潜在机制。

Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B

论文配图:Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B
图 1 · 摘自论文原文
  • 通过注意力探针与消融实验,定位4个关键注意力头在多任务中的核心作用。
  • 在立方体双任务中,模型性能提升59个百分点,显著优于随机初始化模型。
  • 发现跨分布重要性指纹与因果验证一致,适合研究模型可解释性与迁移机制的研究者。

冻结的Gemma 4 31B模型仅在文本数据上预训练,未做修改,通过一个轻量可训练接口迁移到非文本模态。在第24至29层(共192个注意力头)中,基于英文文本的TxtCopy注意力探针(95句)与四个非语言符号任务(二值复制、关联回忆、一维元胞自动机规则90、二进制加法)的逐头消融结果联合识别出4个顶级注意力头:L26.28、L27.28、L27.2、L27.3。该层级联合一致性在超几何零假设下显著(P = 0.0013,N=192,K=38,n=4),并通过多重性校正的置换检验(P_V4 = 0.013)。Gemma L26在OGBench立方体双任务中达60.22%准确率,远超随机初始化模型(约1%),提升59点(n=3);具有正确1/√d_k缩放的冻结随机GPT2对照组亦失败。头级因果验证:在训练好的立方体任务1智能体中,置零L26.28导致成功率从63.3%降至10.0%,而同层低文本复制度负控仅降至46.7%(特异性3.2倍,n=30;配对t检验p=0.039)。全层扫描显示L26.28位列前4/32。诚实负例:同层内文本复制度与消融损失相关系数ρ=+0.37(与层内因果读取相反);单头激活补丁无法传递匹配变量;4个命名头单独不足以完成任一任务;Walker2d-DT与scene-task1任务招募了非命名层(如L24)且无头消融特异性。本研究贡献为跨分布注意力头重要性指纹(层级)及单跨模态目标的头级因果证据。

原文摘要 · Abstract (English)

Frozen Gemma 4 31B weights pretrained exclusively on text, unmodified, transfer through a thin trainable interface to non-text modalities the substrate has never processed. On the L24--L29 slice (192 attention heads), an English-text TxtCopy attention probe (95 sentences) and per-head ablation impact on four non-language token-pattern tasks (binary copy, associative recall, 1D cellular automaton Rule 90, binary addition) jointly classify four heads -- L26.28, L27.28, L27.2, L27.3 -- as top-tier on both signals. The slice-level joint coincidence is significant under hypergeometric null ($P = 0.0013$, $N=192$, $K=38$, $n=4$) and survives multiplicity-aware permutation tests ($P_{V4} = 0.013$). Pretrained Gemma L26 reaches 60.22% on OGBench cube-double-play-task1 vs ~1% for random-init Gemma ($+59$pt at $n=3$); a FrozenRandom-GPT2 control with correct $1/\sqrt{d_k}$ scaling also fails. Head-level causal validation: zeroing L26.28 in the trained cube-task1 IQL agent drops success $63.3\% \to 10.0\%$ vs $46.7\%$ for a layer-matched low-TxtCopy negative control ($3.2\times$ specificity at $n=30$; $n=5$ paired-$t$ $p=0.039$). A full L26 sweep places L26.28 at rank 4 of 32. Honest negatives: within-L26 Spearman $ρ(\text{TxtCopy, drop}) = +0.37$ (opposite of within-layer causal reading); single-head activation patching does not transfer the matching variable; the 4 named heads alone do not suffice on any task; Walker2d-DT and scene-task1 recruit L24 outside the named slice and show null head-ablation specificity. We frame the contribution as a cross-distribution importance fingerprint at the slice level plus head-level causal evidence on one cross-modality target.

注意力头跨模态迁移模型可解释性冻结模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。