发现语言切换可由特定方向激活控制,且存在层间差异。
Steering the Language Axis: From Linear Decodability to Causal Control

- 通过主成分分析提取语言轴,实现对模型输出语言的精准定向控制。
- 在FLORES-200上126万次生成实验中,沿语言轴扰动可稳定触发跨语言切换。
- 语言决策具有层敏感性,且去激活后模型会自动回归英文。
尽管大语言模型具备出色的多语言能力,但其隐含状态中决定语言选择的动态机制仍不清楚。本文探究语言身份是否仅线性可解,或能否通过一个紧凑的激活方向进行因果控制。我们在多个模型家族(包括Qwen 3.5-2B和Llama-3.2-1B-Instruct)中开展全面因果干预分析,基于PCA提取“语言轴”,在FLORES-200数据集上完成126万次生成的转向与消融实验。沿这些几何方向进行扰动可稳定实现跨脚本(英语到中文)与同脚本(英语到西班牙语)的语言切换,而同等幅度的随机扰动几乎无效。分层分析显示,语言承诺高度局部化且依赖具体语言对:英语转中文需后期干预,而英语转西班牙语则更早发生,呈现双峰敏感性。此外,针对性消融实验揭示,一旦语言信号被移除,模型将无一例外回归英语。最终结果表明,语言决策边界在推理过程中表现为方向依赖、层特定的因果活跃特征。
原文摘要 · Abstract (English)
Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood. In this work, we ask whether language identity is merely linearly decodable from hidden states, or if it can be causally controlled by a compact activation direction. We conduct an exhaustive causal intervention analysis across multiple model families, including Qwen 3.5-2B and Llama-3.2-1B-Instruct, isolating PCA-derived "language axes" to perform steering and ablation experiments across 1.26 million generations on the FLORES-200 dataset. Steering along these geometric directions reliably forces language switching in both cross-script (English to Chinese) and same-script (English to Spanish) settings, whereas equal-magnitude random perturbations yield virtually no effect. Our layerwise analysis reveals that language commitment is highly localized and explicitly language-pair-dependent. While English to Chinese switching resists early intervention and steers easily in the later layers, the English-Spanish transition shifts earlier, displaying a distinct, bimodal sensitivity. Furthermore, targeted ablation uncovers a fundamental reversion to English: once the language signal is removed, the model falls back to English regardless of the input prompt. Ultimately, these findings demonstrate that language decision boundaries function during inference as causally active features that are direction-dependent and layer-specific.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。