arXiv:2604.09839cs.AIcs.LG2026-04被引 1

激活操控无法还原为自然文本提示,揭示白盒控制与黑盒提示的本质差异。

Steered LLM Activations are Non-Surjective

论文配图:Steered LLM Activations are Non-Surjective
图 1 · 摘自论文原文
  • 将激活操控的可实现性转化为数学上的满射问题
  • 证明操控后状态无法由任何文本提示复现
  • 警示勿用操控成功误判提示解释力或安全漏洞

激活操控是一种流行的白盒控制技术,通过修改模型内部激活来改变其行为。该技术广泛用于可解释性(如探测真实性)和安全研究(如越狱可行性)。然而,当前尚不清楚这种操控后的状态是否可通过任意文本提示实现。本文将此问题形式化为满射性问题:对于固定模型,每个被操控的激活是否存在一个前像(即自然输入)?在合理假设下,我们证明激活操控会将残差流推向离散提示无法到达的状态空间。几乎必然地,没有任何提示能再现操控带来的内部行为。我们在三个主流大语言模型上实证验证了这一结论。结果建立了白盒可控性与黑盒提示之间的形式化分离。因此,我们警告不应将激活操控的成功视为提示解释力或脆弱性的证据,并主张评估协议应明确区分白盒与黑盒干预。

原文摘要 · Abstract (English)

Activation steering is a popular white-box control technique that modifies model activations to elicit an abstract change in its behavior. It has also become a standard tool in interpretability (e.g., probing truthfulness, or translating activations into human-readable explanations) and safety research (e.g., jailbreakability). However, it is unclear whether steered behavior is realizable by any textual prompt. In this work, we cast this question as a surjectivity problem: for a fixed model, does every steered activation admit a preimage under the model's natural forward pass? Under practical assumptions, we prove that activation steering pushes the residual stream off the manifold of states reachable from discrete prompts. Almost surely, no prompt can reproduce the same internal behavior induced by steering. We also illustrate this finding empirically across three widely used LLMs. Our results establish a formal separation between white-box steerability and black-box prompting. We therefore caution against interpreting the ease and success of activation steering as evidence of prompt-based interpretability or vulnerability, and argue for evaluation protocols that explicitly decouple white-box and black-box interventions.

大模型解释激活操控安全分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。