arXiv:2502.03032cs.LGcs.CL2025-02ICML被引 11

通过追踪特征流动,实现大模型行为的精准可解释与控制。

Analyze Feature Flow to Enhance Interpretation and Steering in Language Models

  • 用无数据余弦相似度追踪特征在层间的演化路径。
  • 发现特征在不同层中持续、转变或首次出现的规律。
  • 可直接放大或抑制特征,实现文本主题的精准控制。

我们提出一种系统化方法,用于追踪稀疏自编码器在大型语言模型连续层间发现的特征,拓展了此前关于层间特征关联的研究。通过无数据余弦相似度技术,我们追踪特定特征在各阶段的持续性、转换或首次出现情况。该方法生成细粒度的特征演化图谱,实现更精细的可解释性与机制洞察。关键的是,我们证明这些跨层特征映射可直接用于调控模型行为,通过增强或抑制特定特征,实现文本生成中的目标主题控制。研究结果表明,因果性跨层可解释框架不仅揭示了特征在前向传播中的发展过程,还为大型语言模型的透明操控提供了新途径。

原文摘要 · Abstract (English)

We introduce a new approach to systematically map features discovered by sparse autoencoder across consecutive layers of large language models, extending earlier work that examined inter-layer feature links. By using a data-free cosine similarity technique, we trace how specific features persist, transform, or first appear at each stage. This method yields granular flow graphs of feature evolution, enabling fine-grained interpretability and mechanistic insights into model computations. Crucially, we demonstrate how these cross-layer feature maps facilitate direct steering of model behavior by amplifying or suppressing chosen features, achieving targeted thematic control in text generation. Together, our findings highlight the utility of a causal, cross-layer interpretability framework that not only clarifies how features develop through forward passes but also provides new means for transparent manipulation of large language models.

可解释性特征追踪模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。