arXiv:2608.28806cs.AI2026-08

通过相似特征分组提升语言模型可控生成效果

Enhancing SAE-based Steering via Neighbor Integrated Feature Selection

论文配图:Enhancing SAE-based Steering via Neighbor Integrated Feature Selection
图 1 · 摘自论文原文
  • 基于特征表示相似性,整合邻近特征信息优化选择
  • 在多个任务中显著优于传统Top-k筛选方法
  • 适合需要精准控制大模型输出的研究者

稀疏自编码器(SAEs)能将模型激活分解为可解释特征,广泛用于大型语言模型的可控生成。现有方法通常基于统计得分选取前k个特征,假设得分高的特征有更强的调控能力。本文发现这一假设常不成立:有效调控特征可能分布在语义相近的相邻特征组中,尽管得分差异大,实际调控效果却相近,导致基于得分的选择会遗漏关键特征。为此,我们提出邻近特征集成选择(NIFS),一种即插即用策略,利用特征表示相似性提升选择精度。我们在多种SAE驱动的调控方法和任务上评估NIFS,结果表明其性能持续优于传统的Top-k选择。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) disentangle model activations into interpretable features and are widely used for steering large language models. Most existing SAE-based steering methods select features by applying a top- filter based on statistical scores, assuming that higher-scoring features yield stronger steering effects. In this paper, we show that this assumption is often invalid, leading to suboptimal feature selection. Our analysis reveals that effective steering features may be distributed among representationally adjacent, semantically similar groups induced by feature splitting in SAEs. Within such groups, features may exhibit disparate statistical scores despite having comparable steering influence, causing score-based selection to overlook important features. Based on these observations, we propose \textsc{Neighbor Integrated Feature Selection} (\textsc{NIFS}), a plug-and-play strategy that leverages representation similarity to improve feature selection for steering. We evaluate \textsc{NIFS} across multiple SAE-based steering methods and tasks, and demonstrate consistent performance gains over conventional top-$k$ selection.

可控生成特征选择语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。