arXiv:2605.29971cs.CL2026-05

通过因果干预改变语言模型中的动词倾向,验证其对语法结构选择的影响。

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

论文配图:Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning
图 1 · 摘自论文原文
  • 在激活向量中定位动词倾向的低维方向,实现连续变量的可编辑干预。
  • 修改动词倾向后,下游句法偏好系统性改变,证明其因果作用。
  • 发现转向向量含误差信号但未被下游任务直接使用,提示学习机制仍存谜题。

语言模型中的因果干预传统上针对离散特征(如数的变化),但语言模型还需处理连续特征。本文提出一种对连续变量进行因果干预的方法:给定与分级目标变量配对的激活向量,我们定位该变量的低维方向,并利用该方向对向量进行反事实编辑。将此方法应用于心理语言学中研究充分的连续特征——动词倾向(反映特定动词后常出现的句法结构),结果表明动词倾向在大语言模型的转向向量中具有因果表征:对动词倾向的反事实修改会系统性改变下游句法偏好。此前研究已发现动词倾向与上下文学习相关;进一步分析显示,转向向量编码了可能驱动错误驱动更新的误差信号,但这些成分在下游生成中并未被因果使用。整体结果表明,因果干预可扩展至连续变量,但连续变量与上下文学习之间的联系仍待揭示。

原文摘要 · Abstract (English)

Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must also make use of features that are graded. We introduce a method for causal intervention on continuous variables: given activation vectors paired with a graded target variable, we localize a low-dimensional direction for that variable and use this direction to edit a vectors toward counterfactual target values. We apply this method to a continuous feature that is well-studied in psycholinguistics, namely verb bias (which reflects which syntactic structures tend to follow a given verb). We show that verb bias is causally represented in steering vectors extracted from large language models: counterfactual edits to verb bias systematically shift downstream structural preferences. Verb bias has also previously been linked to in-context learning; in further analyses, we find that steering vectors encode error signals that could drive the error-driven update behavior seen in in-context learning but that these aspects of the steering vectors are not causally used in downstream production. Overall, these results show causal interventions can be applied to continuous variables, though connecting continuous variables to in-context learning remains a challenge.

因果干预动词倾向上下文学习转向向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。