arXiv:2606.08496cs.CLcs.LG2026-06被引 1

用激活得分引导优化,让SAE特征解释更准更可信

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

论文配图:SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
图 1 · 摘自论文原文
  • 用激活分数作奖励信号,实现解释的自我修正与迭代优化
  • 在因果触发和区分性激活上显著优于现有方法
  • 适合需要可解释、低幻觉的LLM分析场景

尽管稀疏自编码器(SAEs)通过将密集表示分解为稀疏特征缓解了大语言模型(LLMs)的黑箱问题,但解释这些特征仍是核心挑战。现有解释方法多采用开环范式,无法利用机制反馈进行持续优化。本文提出SAEExplainer,一种利用激活得分作为目标奖励信号的训练框架,使模型具备自我纠错与迭代自增强能力。通过两轮优化过程,反复验证并修正基础解释,显著减少解释幻觉,强化因果触发模式。大量实验表明,该方法在多数指标上超越现有基线,尤其在因果触发与区分性激活方面表现突出。

原文摘要 · Abstract (English)

Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features still remains a central challenge. Current explanation methods, however, typically operate within an open-loop paradigm, failing to leverage mechanistic feedback for further refinement. In this paper, we propose SAEExplainer, a training framework utilizes activation scores as an objective reward signal to train the model for self-correction and iterative bootstrapping. By iteratively verifying and correcting foundational explanations through a two-round optimization process, SAEExplainer achieves continuous improvement in its explanatory capabilities. This mechanism significantly reduces explanation hallucinations and reinforces causal triggering patterns. Extensive experiments demonstrate our approach improves upon established baselines across most metrics, especially in causal triggering and discriminative activation.

可解释性SAE自编码器大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。