arXiv:2511.00617cs.LGcs.AI2025-11被引 20

用贝叶斯框架统一解释提示和激活控制,揭示行为突变机制

Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering

  • 将提示与激活干预视为改变模型对潜在概念的信念
  • 预测干预在对数信念空间中具有可加性,导致行为突变
  • 适合研究大模型可控性、推理机制与安全控制的学者

大型语言模型可通过提示(上下文学习)和内部激活(激活引导)在推理时进行控制。尽管已有不同理论解释这些方法,但它们都旨在调控模型行为,这引发了一个问题:这些看似不同的方法是否可被纳入更广泛的统一框架?为此,我们从贝叶斯视角提出一个统一且可预测的模型控制理论。具体而言,我们假设上下文和激活干预均通过改变模型对潜在概念的信念来影响行为:激活引导改变概念先验,而上下文学习则累积证据。由此得到闭式贝叶斯模型,在多个受先前多示例上下文学习启发的领域中高度预测模型行为。该模型不仅能解释已有现象(如提示积累导致的S形学习曲线),还能预测新现象(如两类干预在对数信念空间中可加,引发显著行为跃迁)。本工作为提示与激活控制提供了统一解释,并提供了一种实证预测干预效果的方法。

原文摘要 · Abstract (English)

Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether these seemingly disparate methodologies can be seen as specific instances of a broader framework. Motivated by this, we develop a unifying, predictive account of LLM control from a Bayesian perspective. Specifically, we posit that both context- and activation-based interventions impact model behavior by altering its belief in latent concepts: steering operates by changing concept priors, while in-context learning leads to an accumulation of evidence. This results in a closed-form Bayesian model that is highly predictive of LLM behavior across context- and activation-based interventions in a set of domains inspired by prior work on many-shot in-context learning. This model helps us explain prior empirical phenomena - e.g., sigmoidal learning curves as in-context evidence accumulates - while predicting novel ones - e.g., additivity of both interventions in log-belief space, which results in distinct phases such that sudden and dramatic behavioral shifts can be induced by slightly changing intervention controls. Taken together, this work offers a unified account of prompt-based and activation-based control of LLM behavior, and a methodology for empirically predicting the effects of these interventions.

大模型控制贝叶斯推理上下文学习行为突变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。