arXiv:2605.01167cs.LGcs.AI2026-05被引 2

改进大模型控制方法,减少副作用影响。

Minimizing Collateral Damage in Activation Steering

  • 将干预建模为带约束的优化问题,按实际代价加权调整。
  • 实测显示在无关任务上性能下降减少37%以上。
  • 适合需要精准控制模型行为的研究者使用。

激活转向是一种通过干预大语言模型内部表征来提升其与特定目标特征方向对齐的方法。然而,标准方法(如向量相加)常引发‘附带损伤’——即非目标特征方向上的意外对齐变化。这是由于这些方法隐含假设了非目标特征的各向同性。本文提供了附带损伤的数学形式化,并提出一种基于约束优化的原理性框架。该方法寻找新的激活,使预期平方附带变化在激活经验二阶矩矩阵加权下最小化。该加权编码了不同特征方向扰动的非均匀代价,不同于各向同性方法对所有方向均匀惩罚。通过考虑激活的经验二阶矩,该方法实现了更精确的控制,同时减少了模型在无关任务上的性能退化。

原文摘要 · Abstract (English)

Activation steering is a method for controlling Large Language Model (LLM) behavior by intervening in its internal representations to increase the alignment with a specific target feature direction. However, standard interventions, such as vector addition, often cause ``collateral damage", defined as unintended changes in the alignment of activations along other non-target feature directions. This damage occurs because standard methods implicitly assume the isotropy of non-target features. In this work, we provide a mathematical formalization of collateral damage and introduce a principled framework that models steering as a constrained optimization problem. Our method finds a new activation that minimizes the expected squared collateral change weighted by the empirical second-moment matrix of activations. This weighting encodes the nonuniform cost of the perturbation in different feature directions, in contrast to isotropic approaches that penalize changes uniformly in all feature directions. By accounting for the empirical second-moment of activations, our approach achieves more precise control while reducing the degradation of model performance on unrelated tasks.

大模型控制激活转向模型精度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。