提出用干预评估统一可解释性与控制能力,发现现有方法在控制上效果不稳。
Towards Unifying Interpretability and Control: Evaluation via Intervention
- 将四种可解释方法统一为可干预的编码-解码框架
- 发现基于镜头的方法在简单干预中表现优于稀疏自编码器和探测器
- 机制干预常损害模型连贯性,不如提示词有效
随着大语言模型复杂性和能力的提升,理解其推理过程的需求日益迫切,这通常源于对模型控制与对齐的潜在目标。尽管已有大量可解释性和可控性方法被提出,但它们往往只服务于单一目标,很少兼顾两者。此外,缺乏标准化的应用场景、动机和评估指标,使得难以衡量方法的实际效用。为此,我们主张干预是可解释性的根本目标,并引入评估标准来衡量方法通过干预控制模型行为的能力。我们将四种流行可解释方法——稀疏自编码器(SAE)、logit lens、tuned lens 和探测(probing)——统一并扩展为一个抽象的编码-解码框架,使可解释特征能够被干预并映射回隐层表示以控制输出。我们提出了两个新评估指标:干预成功率和一致性-干预权衡,用于衡量解释的准确性及其在控制中的实用性。研究发现:(1) 当前方法虽支持干预,但其效果在不同特征和模型间不一致;(2) 基于镜头的方法在实现简单具体干预方面优于 SAE 与探测器;(3) 机制干预常导致模型连贯性下降,表现不及更简单的提示策略,暴露出当前可解释方法在需控制的应用中的关键缺陷。
原文摘要 · Abstract (English)
With the growing complexity and capability of large language models, a need to understand model reasoning has emerged, often motivated by an underlying goal of controlling and aligning models. While numerous interpretability and steering methods have been proposed as solutions, they are typically designed either for understanding or for control, seldom addressing both. Additionally, the lack of standardized applications, motivations, and evaluation metrics makes it difficult to assess methods' practical utility and efficacy. To address the aforementioned issues, we argue that intervention is a fundamental goal of interpretability and introduce success criteria to evaluate how well methods can control model behavior through interventions. To evaluate existing methods for this ability, we unify and extend four popular interpretability methods-sparse autoencoders, logit lens, tuned lens, and probing-into an abstract encoder-decoder framework, enabling interventions on interpretable features that can be mapped back to latent representations to control model outputs. We introduce two new evaluation metrics: intervention success rate and coherence-intervention tradeoff, designed to measure the accuracy of explanations and their utility in controlling model behavior. Our findings reveal that (1) while current methods allow for intervention, their effectiveness is inconsistent across features and models, (2) lens-based methods outperform SAEs and probes in achieving simple, concrete interventions, and (3) mechanistic interventions often compromise model coherence, underperforming simpler alternatives, such as prompting, and highlighting a critical shortcoming of current interpretability approaches in applications requiring control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。