发现Transformer将语义概念编码在低方差谱尾方向,形成双重几何结构。
Concepts Whisper: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations

- 通过差异均值向量、SAE特征与线性探测,发现概念方向在谱尾反集中。
- 17个模型中13个保持谱尾对齐,语法信息更倾向高方差子空间。
- 适合研究模型内部表征机制或干预设计的读者关注。
我们发现,变换器中的概念表征在未嵌入协方差的谱尾系统性地反集中,将词级概念编码于17个核心模型和22个语义类别中的低方差方向。三类独立提取方法均得收敛结果:残差流差异均值向量在全部17个模型中反集中(单样本t检验,p = 3.8e-9),且在13/17个模型中仍比范数匹配的随机方向更接近谱尾。稀疏自编码器(SAE)特征在模型内跨概念呈现显著反集中(p = 4.5e-19)。线性探测在Llama与Qwen上亦支持该现象。研究揭示双重几何:激活空间的概念方向反集中,而静态未嵌入行对比则集中在高方差方向(p < 10^-4)。该研究源于验证Park等人(2024)因果内积是否促进跨语言概念迁移;在17个模型与4种语言对上的匹配谱随机化实验显示,白化因果对齐并未优于谱正则化本身(p = 0.95)。分裂注入干预在五种模型中四次展现预测的干扰不对称性(配对Cohen's d_z最高达1.19),该范围内无显著反转;八个模型的词性标注探测表明,六种架构中语法偏好高方差子空间,但在Qwen 2.5家族中出现显著反转。结果提示:变换器在上下文处理中将语义内容旋转至谱静区,某些架构中,低方差引导可减少语法破坏。
原文摘要 · Abstract (English)
We find that transformer concept representations systematically anti-concentrate in the spectral tail of the unembedding covariance, encoding word-level concepts in low-variance directions across a 17-model core suite and an expanded set of 22 semantic concept categories, with convergent replications from three independent extraction methods. Residual-stream difference-of-means vectors anti-concentrate in all 17 models (model-level one-sample t-test, p = 3.8e-9), and remain more tail-aligned than norm-matched random directions in 13 of 17; convergent support comes from sparse autoencoder (SAE) features (p = 4.5e-19 across concepts within a model) and linear probes on Llama and Qwen. We identify a dual geometry: activation-space concept directions anti-concentrate while static unembedding-row contrasts concentrate in high-variance directions (p < 10^-4). This investigation arose from testing whether the causal inner product of Park et al. (2024) aids cross-lingual concept transport; a matched-spectrum randomization across 17 models and four language pairs finds no evidence that Whitened Causal Alignment improves over spectral regularization alone (p = 0.95). Split-injection interventions, restricted to steering strengths at which both arms remain interpretable, show the predicted interference asymmetry in four of five models (paired Cohen's d_z up to 1.19) with no significant reversal inside that regime, and POS-tag probing across eight models shows syntax preferentially encoded in the high-variance subspace in six of eight architectures, with a significant reversal in the Qwen 2.5 family. These results suggest transformers rotate semantic content into spectrally quiet regions during contextualized processing, where, in some architectures, interventions may reduce grammatical disruption relative to high-variance steering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。