编辑实验难以证明大模型行为被局部化
Does Editing Provide Evidence for Localization?
- 用对齐技术寻找最优局部编辑方案
- 随机位置的编辑效果与全局对齐相当
- 编辑有效不等于该位置编码目标行为
大型语言模型可解释性研究的一个基本目标是将语义行为定位到模型内部特定组件。现有方法通过启发式手段寻找候选定位区域,再通过修改对应内部表征并观察模型行为是否符合语义解释来评估定位有效性。本文探讨此类编辑证据的强度。我们提出一种新方法,利用大模型对齐技术寻找最优局部编辑。实验显示,尽管某处编辑看似强烈支持定位假设,但实际定位完全失败;且在随机位置进行最优编辑,效果可媲美全模型对齐。综合结果表明,仅凭局部编辑能引发目标行为改变,无法提供可靠证据证明该位置确实编码了该行为。
原文摘要 · Abstract (English)
A basic aspiration for interpretability research in large language models is to "localize" semantically meaningful behaviors to particular components within the LLM. There are various heuristics for finding candidate locations within the LLM. Once a candidate localization is found, it can be assessed by editing the internal representations at the corresponding localization and checking whether this induces model behavior that is consistent with the semantic interpretation of the localization. The question we address here is: how strong is the evidence provided by such edits? To evaluate the localization claim, we want to assess the effect of the optimal intervention at a particular location. The key new technical tool is a way of adapting LLM alignment techniques to find such optimal localized edits. With this tool in hand, we give an example where the edit-based evidence for localization appears strong, but where localization clearly fails. Indeed, we find that optimal edits at random localizations can be as effective as aligning the full model. In aggregate, our results suggest that merely observing that localized edits induce targeted changes in behavior provides little to no evidence that these locations actually encode the target behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。