arXiv:2602.00945cs.CLcs.AI2026-02

通过操控特定神经元,让大模型优先用印地语或西班牙语回答。

Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs

  • 定位语言专属神经元,识别影响语言偏好的关键激活单元。
  • 发现低秩控制空间,仅用少量方向就能改变模型语言默认偏好。
  • 可精准调整模型语言倾向,适合多语言场景下的个性化优化。

大型语言模型虽具备多语言能力,但预训练数据以英语为主,导致其他语言在参数记忆中被系统性抑制。本文提出神经FOXP2方法,通过操控语言特异性神经元,使模型以印地语或西班牙语为默认语言。该方法分三步:(i) 局部化——每层训练稀疏自编码器(SAE),将激活分解为少量特征组件,量化每个特征对目标语言的偏好程度,追踪最强贡献单元,形成紧凑的语言神经元集合;(ii) 驱动方向——基于各层英语与目标语言间的激活差矩阵,进行分层奇异值分解(SVD),提取主导变化方向,利用特征值间隙和有效秩谱确定可控干预窗口;(iii) 驱动——在低至中层施加有符号的稀疏激活偏移:沿目标语言主方向正向调整,同时对英语神经元施加反向补偿,实现可控的语言默认偏好切换。

原文摘要 · Abstract (English)

LLMs are multilingual by training, yet their lingua franca is often English, reflecting English language dominance in pretraining. Other languages remain in parametric memory but are systematically suppressed. We argue that language defaultness is governed by a sparse, low-rank control circuit, language neurons, that can be mechanistically isolated and safely steered. We introduce Neural FOXP2, that makes a chosen language (Hindi or Spanish) primary in a model by steering language-specific neurons. Neural FOXP2 proceeds in three stages: (i) Localize: We train per-layer SAEs so each activation decomposes into a small set of active feature components. For every feature, we quantify English vs. Hindi/Spanish selectivity overall logit-mass lift toward the target-language token set. Tracing the top-ranked features back to their strongest contributing units yields a compact language-neuron set. (ii) Steering directions: We localize controllable language-shift geometry via a spectral low-rank analysis. For each layer, we build English to target activation-difference matrices and perform layerwise SVD to extract the dominant singular directions governing language change. The eigengap and effective-rank spectra identify a compact steering subspace and an empirically chosen intervention window (where these directions are strongest and most stable). (iii) Steer: We apply a signed, sparse activation shift targeted to the language neurons. Concretely, within low to mid layers we add a positive steering along the target-language dominant directions and a compensating negative shift toward the null space for the English neurons, yielding controllable target-language defaultness.

语言控制神经元操控多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。