提出单神经元控制的预算窗口定律,预测何时能有效操控模型行为而不崩溃。
Leverage Is Not Reach: A Control-Window Law for Single-Neuron Steering in Language Models

- 用残差流与写入方向的对齐度作为控制坐标,受统一饱和曲线约束。
- 在15个神经元上预测崩溃阈值误差仅0.14,11个案例优于多数基线。
- 揭示控制失败三类原因,解释为何梯度归因常误判控制器。
对齐语言模型通过稀疏前馈神经元控制拒绝行为和语言路由,但尚无理论可预测单神经元干预何时能实现一致控制而非输出崩溃。本文提出一种归一化预算的控制窗口框架:沿单一写入方向施加剂量后,控制效果由残差流与写入方向的对齐度决定,遵循由残差范数与写入范数比值设定的通用饱和曲线。当行为触发点低于崩溃上限时,控制才保持一致。该同一坐标同时支配良性模式切换与拒绝行为;上限由权重和一次前向传播确定,而触发点在推理过程中测量。在15个保留神经元上,预测上限平均绝对误差为0.14(批量层约0.07),所作闭合判断在11个案例中胜过15个中的10个多数基线。闭合案例揭示三种失败模式而非违反规律:触发前崩溃、深度不足无法传播或归一化限制推动距离。该定律解释了为何局部梯度归因反向预测控制——真正控制器沿读出轴写入,其一阶梯度接近零。仅依赖前向传播的对比筛选机制经窗口精确化后,可恢复归因遗漏的控制器。在拒绝任务中,干预成功具有类型差异:流畅且非行动内容文本中可实现协同绕过,而真正行动可达性仅出现在六组审计的Llama枢纽中三个,并需在后期推理阶段出现。因此,单神经元控制本质上是受预算约束的类型化可操控性审计,而非固定剂量的轶事。
原文摘要 · Abstract (English)
Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collapsing the output. We develop a budget normalized control window framework for single neuron steering. A dose along one write direction reduces to one control coordinate: the alignment between the residual stream and the write, driven along a universal saturation curve in units of a coherence budget set by the residual norm divided by the write norm. Coherent control exists when a behavior trigger lies below the collapse ceiling. The same coordinate governs benign mode switches and refusal; the ceiling follows from weights and one generic forward pass, while triggers are measured at rollout. On fifteen held out neurons, the predicted ceiling has mean absolute error 0.14, about 0.07 in bulk layers, and the committed open or closed verdict holds on eleven against a ten of fifteen majority baseline. Closed cases expose three failure modes rather than violations: collapse before trigger, too little depth to propagate, or a normalization that caps how far one neuron can push. The law explains why local gradient attribution anti predicts control: true controllers write off the readout axis and carry a near zero first order gradient. A forward only contrastive screen made precise by the window recovers controllers that attribution misses. On refusal, the hardest case, intervention success is typed, not scalar: coherent bypass and strict actionable reach separate, so a neuron can flip refusal in fluent, on task text with no actionable content, and genuine actionable reach appears only for three of six audited Llama pivots and only at later rollout horizons. Single neuron steering is therefore a budgeted, typed audit of controllability rather than a fixed dose anecdote.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。