提前预测稀疏自编码器干预的副作用,提升语言模型控制的精准性。
Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects

- 基于激活统计与解码器几何结构,预判特征干预的模块化程度。
- 在GPT-2-small等模型中,预测信号强且能抵抗残差干扰。
- 不同模型需关注不同副作用维度,如稳定性或无关扰动范围。
稀疏自编码器(SAE)特征被广泛用于控制语言模型,但干预常不干净:同一操作在不同上下文中表现不一,且会扰动无关特征。本文提出一种干预前筛选框架,通过干预前的特征统计量预测SAE干预的副作用。从干预模块性、效果稳定性和旁侧扩散三个维度评估GPT-2-small、Pythia-70M-deduped、Gemma-2-2B和Llama-3.1-8B在ReLU、JumpReLU和TopK SAE字典下的表现。结果表明,解码器几何、激活统计、共激活结构及直接对数项足迹的组合,优于仅频率或激活幅值的基线。该信号在GPT-2-small、Pythia-70M和Llama-3.1-8B中最强,可抵抗幅度相关混杂因素;而在Gemma-2-2B中较弱。留出测试显示,按预测清洁度排序未见特征可选出在新上下文更干净的干预项,但有效轴因模型而异:GPT-2主要提升整体清洁度,Pythia侧重稳定性,Llama聚焦旁侧扩散,Gemma仅部分有效。控制变量实验表明,当字典宽度从32K增至128K时,预测信号仍存在,但收益稳定性下降。总体而言,SAE副作用可预先预测,但有效预测信号和转移的模块性维度依赖于具体模型与字典设置。
原文摘要 · Abstract (English)
Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across contexts and perturb unrelated features. We introduce a pre-intervention screening framework for forecasting SAE steering side effects from feature statistics computed before steering. We operationalize side effects along two axes of steering modularity, effect stability and collateral spread, and evaluate GPT-2-small, Pythia-70M-deduped, Gemma-2-2B, and Llama-3.1-8B across ReLU, JumpReLU, and TopK SAE dictionaries. Across these settings, decoder geometry, activation statistics, co-activation structure, and direct-logit footprint predict steering modularity better than frequency-only and activation-magnitude baselines. The signal is strongest in GPT-2-small, Pythia-70M, and Llama-3.1-8B, where it survives residualization against magnitude-related confounds, and weaker in Gemma-2-2B. Held-out screening shows that ranking unseen features by predicted cleanliness can select features that steer more cleanly on fresh contexts, but the successful axis varies by setting: GPT-2 improves most cleanly, Pythia improves mainly on stability, Llama mainly on collateral, and Gemma only partially. A controlled Llama Scope width comparison shows that the predictive signal persists under a 32K-to-128K dictionary-width change, although the screening payoff becomes less stable. Overall, SAE steering side effects are predictable in advance, but the useful predictor signature and transferred modularity axis are model- and dictionary-setting dependent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。