突破线性干预局限,实现对大模型非线性特征的精准操控
Non-linear Interventions on Large Language Models
- 提出非线性干预通用框架,可作用于隐含在非线性流形上的特征
- 在拒绝规避控制任务中,通过干预非线性拒答特征,精度显著优于线性基线
- 适用于需要精细控制模型行为的场景,如安全对齐与可控生成
干预是理解大语言模型内部表示的重要且广泛应用的方法。然而,现有干预方法受限于基于线性表示假设的线性干预,无法触及编码在非线性流形上的特征。本文提出一种自然扩展至非线性特征表示的通用干预形式,并设计相应学习机制,使干预可作用于无直接输出信号的隐式特征。我们在拒绝规避引导任务中验证该框架,结果表明,通过干预一个控制拒绝行为的非线性特征,该方法比线性基线更精确地引导模型行为。
原文摘要 · Abstract (English)
Intervention is one of the most representative and widely used methods for understanding the internal representations of large language models (LLMs). However, existing intervention methods are confined to linear interventions grounded in the Linear Representation Hypothesis, leaving features encoded along non-linear manifolds beyond their reach. In this work, we introduce a general formulation of intervention that extends naturally to non-linearly represented features, together with a learning procedure that further enables intervention on implicit features lacking a direct output signature. We validate our framework on refusal bypass steering, where it steers the model more precisely than linear baselines by intervening on a non-linear feature governing refusal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。