arXiv:2505.20063cs.LGcs.AI2025-05EMNLP被引 71

筛选高质量特征可让无监督模型控制更有效,效果媲美有监督方法。

SAEs Are Good for Steering -- If You Select the Right Features

  • 区分输入型与输出型特征,用双指标评估其价值。
  • 过滤低输出分数特征后,控制效果提升2-3倍。
  • 适合希望低成本实现精准模型控制的研究者。

稀疏自编码器(SAEs)被提出作为无监督分解模型隐空间的方法,可用于无需标注数据的模型控制(steering)。现有方法通过分析激活特征的输入词元来识别可用特征,但近期研究指出,仅凭激活无法完整描述特征对模型输出的影响。本文区分两类特征:输入特征主要捕捉输入模式,输出特征则对模型输出具有人类可理解的影响。我们提出输入与输出评分以表征和定位这两类特征,发现两者高分极少共现。实践表明,剔除低输出评分特征后,使用SAE进行控制的效果提升2-3倍,使其性能接近有监督方法。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have been proposed as an unsupervised approach to learn a decomposition of a model's latent space. This enables useful applications such as steering - influencing the output of a model towards a desired concept - without requiring labeled data. Current methods identify SAE features to steer by analyzing the input tokens that activate them. However, recent work has highlighted that activations alone do not fully describe the effect of a feature on the model's output. In this work, we draw a distinction between two types of features: input features, which mainly capture patterns in the model's input, and output features, which have a human-understandable effect on the model's output. We propose input and output scores to characterize and locate these types of features, and show that high values for both scores rarely co-occur in the same features. These findings have practical implications: after filtering out features with low output scores, we obtain 2-3x improvements when steering with SAEs, making them competitive with supervised methods.

模型控制稀疏编码无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。