发现大模型拒绝行为的路由机制,可精准控制安全响应。
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models

- 通过注意力门控与放大头构成路由电路,决定模型是否拒绝
- 720亿参数模型中单个头影响强度达58倍,门控是必要组件
- 改变检测层信号可实现从拒答到有害引导的连续控制
我们定位了对齐训练语言模型中的策略路由机制。中间层注意力门控读取检测内容后,触发深层放大头增强拒绝信号。在小模型中,该门控与放大头为单个注意力头;在大模型中则扩展为相邻层上的头组。门控贡献输出差异量(DLA)不足1%,但互换测试(p < 0.001)和敲除级联实验证明其因果必要性。在≥120个模型中,跨六家实验室(20亿至720亿参数)均检测到相同模式,具体头因实验室而异。单头消融最多削弱58倍,且会遗漏互换识别出的门控;在大规模下,互换是唯一可靠审计方法。连续调节检测层信号可实现从强硬拒绝、规避到事实回答的政策调控。在安全提示下,同一干预使拒绝转为有害指导,表明安全能力由路由控制而非移除。阈值随主题与输入语言变化,电路在同家族生成间迁移,但行为基准无变。路由为早期承诺:门控在自身层即触发,早于深层处理完成。上下文替换密码使门控互换必要性降低70%~99%,模型转向解谜而非拒绝;将明文门控激活注入密码前向传播,可恢复Phi-4-mini 48%的拒绝率,定位绕过点至路由接口。第二种方法——密码对比分析,利用明文/密文DLA差异,在O(3n)前向传递内映射全敏感路由电路。任何破坏检测层模式匹配的编码均能绕过策略,无论深层是否重建内容。
原文摘要 · Abstract (English)
We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate and amplifier are single heads; at larger scale they become bands of heads across adjacent layers. The gate contributes under 1% of output DLA, yet interchange testing (p < 0.001) and knockout cascade confirm it is causally necessary. Interchange screening at n >= 120 detects the same motif in twelve models from six labs (2B to 72B), though specific heads differ by lab. Per-head ablation weakens up to 58x at 72B and misses gates that interchange identifies; at scale, interchange is the only reliable audit. Modulating the detection-layer signal continuously controls policy from hard refusal through evasion to factual answering. On safety prompts the same intervention turns refusal into harmful guidance, showing that the safety-trained capability is gated by routing, not removed. Thresholds vary by topic and by input language, and the circuit relocates across generations within a family even while behavioral benchmarks register no change. Routing is early-commitment: the gate fires at its own layer before deeper layers finish processing the input. An in-context substitution cipher collapses gate interchange necessity by 70 to 99% across three models, and the model switches to puzzle-solving rather than refusal. Injecting the plaintext gate activation into the cipher forward pass restores 48% of refusals in Phi-4-mini, localizing the bypass to the routing interface. A second method, cipher contrast analysis, uses plain/cipher DLA differences to map the full cipher-sensitive routing circuit in O(3n) forward passes. Any encoding that defeats detection-layer pattern matching bypasses the policy regardless of whether deeper layers reconstruct the content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。