arXiv:2602.16080cs.CLcs.CY2026-02被引 5

通过生成因果中介定位并控制语言模型长文本中的扩散行为

Activation Steering via Generative Causal Mediation

  • 构建对比输入与长文本响应数据集,识别关键注意力头
  • 用稀疏注意力头控制拒绝、奉承、风格转换等行为,效果优于相关方法
  • 适用于需要精准干预长文本行为的研究者与开发者

在长文本回答中,如何定位并控制分散于多个词元的行为?我们提出生成因果中介(GCM)方法,通过对比长文本响应数据集,筛选出能中介特定概念的模型组件(如注意力头),实现对扩散性概念(如诗歌式表达与散文式表达)的精准控制。GCM首先构建包含对比行为输入和长文本输出的数据集,量化各组件对概念的中介作用,选择最强中介组件进行干预。我们在三种语言模型上评估了拒绝、奉承和风格迁移三种行为,结果表明,GCM能有效定位长文本中的概念,并在仅使用少量注意力头的情况下,性能超越基于相关性探测的基线方法。这些结果证明,GCM为语言模型长文本行为的定位与控制提供了有效路径。

原文摘要 · Abstract (English)

Where should we intervene in a language model (LM) to localize and control behaviors that are diffused across many tokens of a long-form response? We introduce Generative Causal Mediation (GCM), a procedure for selecting model components (e.g., attention heads) from contrastive long-form responses, to steer such diffuse concepts (e.g., talk in verse vs. talk in prose). In GCM, we first construct a dataset of contrasting behavioral inputs and long-form responses. Then, we quantify how model components mediate the concept and select the strongest mediators for steering. We evaluate GCM on three behaviors--refusal, sycophancy, and style transfer--across three language models. GCM successfully localizes concepts expressed in long-form responses and outperforms correlational probe-based baselines when steering with a sparse set of attention heads. Together, these results demonstrate that GCM provides an effective approach for localizing from and controlling the long-form responses of LMs.

行为控制因果中介注意力头长文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。