arXiv:2603.00029cs.CL2026-03ACL

发现大模型中关键特征维度可作为可解释的语义控制开关。

Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models

  • 通过激活大小识别出关键功能维度,无需训练即可定位。
  • 仅操纵关键维度就能实现更好领域迁移与越狱攻击效果。
  • 适合想理解或控制大模型内部机制的研究者与工程师。

大型语言模型(LLMs)具有高度各向异性的内部表示,常表现为大量激活——少数特征维度的数值远超其余维度。以往研究将这些极端维度视为需处理的副作用,本文提出新视角:这些维度是领域专业化带来的内在可解释功能单元。我们提出一种基于幅度的简单判别准则,无需训练即可识别出领域关键维度。分析表明,这些维度能作为符号/量化模式或特定领域术语的可解释语义检测器。此外,我们引入关键维度操控(Critical Dimension Steering),仅对识别出的维度施加激活操控。实验证明,该方法在领域适应和越狱场景中优于传统全维度操控。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit highly anisotropic internal representations, often characterized by massive activations, a phenomenon where a small subset of feature dimensions possesses magnitudes significantly larger than the rest. While prior works view these extreme dimensions primarily as artifacts to be managed, we propose a distinct perspective: these dimensions serve as intrinsic interpretable functional units arising from domain specialization. Specifically, we propose a simple magnitude-based criterion to identify Domain-Critical Dimensions in a training-free manner. Our analyses reveal that such dimensions behave as interpretable semantic detectors for symbolic/quantitative patterns or domain-specific terms. In addition, we introduce Critical Dimension Steering, which applies activation steering exclusively to the identified dimensions. Empirical results show that this approach outperforms conventional whole-dimension steering in domain adaptation and jailbreaking scenarios.

大模型解释特征操控可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。