arXiv:2411.08790cs.LGcs.AI2024-11被引 21

探究稀疏自编码器为何无法准确解析控制向量的机制

Can sparse autoencoders be used to decompose and interpret steering vectors?

  • 发现控制向量超出自编码器训练分布范围
  • 指出负向特征投影导致重构失败
  • 揭示直接使用自编码器解释控制向量的局限性

控制向量是调控大语言模型行为的有力手段,但其内在机制仍不清晰。尽管稀疏自编码器(SAEs)可能提供一种解析方法,但近期研究发现,由SAE重构的向量往往失去原始向量的控制能力。本文深入分析了直接将SAE应用于控制向量时产生误导性分解的原因:(1) 控制向量偏离了SAE所针对的输入数据分布;(2) 控制向量在某些特征方向上存在有意义的负投影,而SAE并未为此类情况设计。这些限制阻碍了SAE在解释控制向量中的直接应用。

原文摘要 · Abstract (English)

Steering vectors are a promising approach to control the behaviour of large language models. However, their underlying mechanisms remain poorly understood. While sparse autoencoders (SAEs) may offer a potential method to interpret steering vectors, recent findings show that SAE-reconstructed vectors often lack the steering properties of the original vectors. This paper investigates why directly applying SAEs to steering vectors yields misleading decompositions, identifying two reasons: (1) steering vectors fall outside the input distribution for which SAEs are designed, and (2) steering vectors can have meaningful negative projections in feature directions, which SAEs are not designed to accommodate. These limitations hinder the direct use of SAEs for interpreting steering vectors.

控制向量自编码器可解释性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。