arXiv:2604.02608cs.LG2026-04被引 3

功能向量可操控模型行为,却无法通过词元解码器识别,揭示了模型内部指令编码的隐蔽性。

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

论文配图:Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens
图 1 · 摘自论文原文
  • 用函数向量(FV)在激活空间中线性操控模型行为,验证其是否可被解码
  • 4032组实验显示95%以上成功操控但无法解码,仅3例可解码但不可操控
  • 非线性结构隐藏计算指令,现有解码工具对主流模型干预无效

激活操控假设任务相关行为对应激活空间中的线性方向——既可操控也可通过反嵌入解码。函数向量(FVs)作为ICL演示均值差的典型实例,本研究在12个任务、6个模型(3个模型族)、4032个跨模板对上检验该假设。结果相反:FV操控普遍成功,而词元解码器在任意中间层均无法识别正确答案;反之,可解码但不可操控的情况几乎不存在(仅3例)。该差异非表征方言所致。单对角调谐解码器仅解决14例中的1例;二层MLP探测器结合控制实验解决10例中的5例,仍遗留5例对所有测试解码器完全不可见。即使操控准确率超0.90,投影至反嵌入后仍得无意义词元分布。表明FV编码的是计算指令而非答案方向。模型家族差异显著:Mistral FV重写中间表示,而Llama与Gemma FV操控最终输出却无词元解码痕迹,由三重信号(操控后变化、激活修补恢复、范数转移相关性)证实。此前报道的负余弦转移相关性在大尺度下消失,增量贡献ΔR²≤0.011。研究将线性表示假说拆分为线性可解码与线性可操控,并证明二者分离且违背直觉,对安全监控有重要启示:基于词汇投影的工具对广泛部署模型的此类干预完全失明。

原文摘要 · Abstract (English)

Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both steer the model and be readable along the unembedding. Function vectors (FVs), extracted as mean differences across ICL demonstrations, are the canonical test case; the prediction: steering and decoding succeed or fail together. Across 12 tasks, 6 models from 3 families, and 4,032 directed cross-template pairs, we find the opposite. FV steering routinely succeeds where the logit lens cannot decode the correct answer at any intermediate layer, while the converse -- decodable without steerable -- is nearly empty (3 of 72). The gap is not representational dialect. A diagonal tuned lens closes 1 of 14 steerable-not-decodable cases; a 2-layer MLP probe with a Hewitt \& Liang control closes 5 of 10 via nonlinearly encoded structure but leaves 5 invisible to every decoder tested. Even at $> 0.90$ steering accuracy, projecting the FV through the unembedding yields incoherent token distributions: FVs encode computational instructions, not answer directions. A model-family asymmetry sharpens the picture. Mistral FVs rewrite intermediate representations, while Llama and Gemma FVs steer the final output without leaving a logit-lens-visible trace, corroborated by three signals (post-steering deltas, activation-patching recovery, FV norm-transfer correlations). A previously reported negative cosine-transfer correlation dissolves at scale, adding at most $ΔR^2 = 0.011$ beyond task identity. These results decompose the linear representation hypothesis into linear decodability and linear steerability and show they come apart opposite to intuition, with implications for safety monitoring: vocabulary-projection tools are blind to FV-style interventions on widely deployed model families.

模型可解释性函数向量激活操控安全监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。