提出四元操作框架,统一解释Transformer内部计算机制。
Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
- 用对齐、整合、抑制、路由四类操作描述Transformer计算过程。
- 抑制与路由权重模式在15个模型中显著优于随机基线。
- 适合研究模型可解释性及跨架构对比的学者参考。
机制可解释性已识别出Transformer电路,但缺乏描述其功能在不同任务与架构间如何组合的通用语言。本文提出一致主义概率组合主义(CPC)框架,将Transformer计算建立在解释的一致性理论基础上,并通过四种操作角色进行描述:对齐识别候选关系,整合融合支持信息,抑制排除不兼容选项,路由将选定信息传递至输出。在五个架构家族的15个模型中,抑制、整合与路由的权重空间特征与保留激活水平的角色度量相关性高于随机基线。抑制在任务间更稳定。移除对齐头会降低下游抑制活性,且在10个模型中超出随机头控制,但在无冲突提示下也出现类似效应,表明其为上游依赖而非矛盾特异性耦合。显式矛盾显著改变14个模型的层间一致性代理值;去除共享残差协方差后,所有模型均呈现预测方向差异。基础模型与指令微调版本保持归纳头得分结构(r≥0.98),但操作符特征未系统向深层迁移。结果支持CPC作为比较Transformer机制的共享词汇,同时表明其深度与几何表达仍具架构特异性。
原文摘要 · Abstract (English)
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。