arXiv:2603.09313cs.AI2026-03被引 3

提出非线性干预方法,让大模型行为控制更精准

Curveball Steering: The Right Direction To Steer Isn't Always Linear

  • 用多项式核PCA在特征空间进行非线性干预
  • 在几何扭曲强的场景下表现显著优于线性方法
  • 适合需要精准控制模型行为的研究者

激活值调制是通过干预内部表示来控制大语言模型行为的常用方法。现有方法多基于全局线性假设,认为行为属性可通过线性方向调节。然而实践中,线性干预常表现不一致。我们通过测量测地距离与欧氏距离之比,分析了大模型激活空间的内在几何结构,发现存在显著且依赖概念的几何畸变,表明激活空间无法被全局线性几何良好近似。受此启发,我们提出「Curveball steering」,一种基于多项式核PCA的非线性调制方法,在特征空间中执行干预,更尊重学习到的激活几何结构。该方法在几何畸变强烈的场景下持续优于基于线性PCA的调制,表明几何感知的非线性调制为全局线性干预提供了更合理的替代方案。

原文摘要 · Abstract (English)

Activation steering is a widely used approach for controlling large language model (LLM) behavior by intervening on internal representations. Existing methods largely rely on the Linear Representation Hypothesis, assuming behavioral attributes can be manipulated using global linear directions. In practice, however, such linear interventions often behave inconsistently. We question this assumption by analyzing the intrinsic geometry of LLM activation spaces. Measuring geometric distortion via the ratio of geodesic to Euclidean distances, we observe substantial and concept-dependent distortions, indicating that activation spaces are not well-approximated by a globally linear geometry. Motivated by this, we propose "Curveball steering", a nonlinear steering method based on polynomial kernel PCA that performs interventions in a feature space, better respecting the learned activation geometry. Curveball steering consistently outperforms linear PCA-based steering, particularly in regimes exhibiting strong geometric distortion, suggesting that geometry-aware, nonlinear steering provides a principled alternative to global, linear interventions.

大模型控制非线性干预几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。