发现大模型在特征方向上具备特异性纠错能力,支持超叠加计算机制。
Evidence for feature-specific error correction in LLMs

- 通过扰动激活值测试模型纠错能力,发现特征方向更敏感。
- 数学建模显示 $p>2$ 的响应模式,符合特征特异性纠错预测。
- 多模型验证结果一致,适用于理解大模型内部表征机制的科研人员。
理解大语言模型(LLMs)的特征是可解释性研究的核心目标。通常认为 LLMs 使用超叠加(superposition)来表示超过其维度数量的特征,并可能在超叠加状态下执行计算。理论预测,超叠加计算需要一种偏爱特征方向而非通用方向的误差校正机制,但这一预测尚未得到实证检验。本文提出基于激活扰动的实证测试方法:扰动残差流激活,发现其对小扰动具有鲁棒性,形成激活平台,符合误差校正特征;但在候选特征方向(由对比提示对构造的“纯”方向)上的鲁棒性低于两个此类方向的混合方向,表明特征方向被优先对待。通过将扰动效应建模为特征分量分解的 $L^p$-范数函数,发现当 $p=2$ 时响应为二次型,非零特征值不超过残差流维度,无法支持所需大量特征方向的偏倚;而 $p>2$ 则解除此约束,与特征特异性纠错一致。实验结果在 Gemma-2-9B、Qwen3-1.7B、Llama-3.1-8B、Mistral-7B-v0.3、Aya-Expanse-8B、Yi-1.5-9B 六个模型中复现。进一步在具有已知真实特征的玩具模型上验证,恢复 $p>2$ 于真实特征方向,随偏离角度增加逐渐趋近 $p ightarrow2$。
原文摘要 · Abstract (English)
Understanding the features of large language models (LLMs) is a central goal of interpretability. LLMs are commonly assumed to use superposition to represent more features than they have dimensions. They may not only represent features in superposition but also perform computation in superposition. Theory predicts that computing in superposition requires error correction that privileges feature directions over generic ones, but this prediction has not been tested empirically. We propose an empirical test of error correction in LLMs based on activation perturbations. Perturbing residual-stream activations, we find that they are robust to small perturbations--forming activation plateaus consistent with error correction--but less robust along candidate feature directions ("pure" directions, constructed from contrastive prompt pairs) than along mixtures of two such directions, indicating that the pure directions are privileged. We quantify this privilegedness by modeling the perturbation effect as a function of the $L^p$-norm of its decomposition into feature components. For $p=2$ the response is a quadratic form with at most as many nonzero eigenvalues as the residual-stream dimension, which cannot privilege the many feature directions superposition requires. $p>2$ lifts this constraint and is consistent with feature-specific error correction. We find $p>2$ for contrastive, MELBO, and SAE-decoder directions, and $p\approx2$ for random and PCA directions (controls). These results replicate across Gemma-2-9B, Qwen3-1.7B, Llama-3.1-8B, Mistral-7B-v0.3, Aya-Expanse-8B, and Yi-1.5-9B. We further validate our method on a toy model of error correction with known ground-truth features, recovering $p>2$ for true feature directions, degrading toward $2$ as we rotate away from them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。