让语言模型概念更独立,减少干预时的干扰。
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
- 通过近正交化特征字典,减少概念间的纠缠。
- 在数学推理任务上实现更精准的局部干预。
- 适合研究模型可解释性与因果干预的学者。
机制可解释性的核心假设是语言模型中的有意义概念由激活空间中的线性特征表示。为支持可靠干预,操纵一个特征不应显著影响其他特征的效果。然而,现实中特征纠缠导致干扰,使局部干预产生意外下游影响。受独立因果机制原则启发,我们提出约束内部特征近乎正交。这能促进模块化表征,便于因果干预。我们通过特征干扰的差距来形式化理想隔离干预与实际输出效果之间的差异,并用特征字典的自一致性上界了干扰传播。该偏差与字典上的显式正交性正则化相关。实验表明,该正则化在保持模型性能的同时,提升了数学推理概念的隔离干预能力。代码已开源:https://github.com/mrtzmllr/sae-icm。
原文摘要 · Abstract (English)
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the _Independent Causal Mechanisms_ principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the feature dictionary, and relate this discrepancy to an explicit orthogonality regularization on the dictionary itself. Empirically, we show that this regularization enables more isolated interventions on mathematical reasoning concepts while preserving model performance. Our code is available under https://github.com/mrtzmllr/sae-icm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。