用线性回归高效学习大模型内部可解释的稀疏电路
Scalable Circuit Learning for Interpreting Large Language Models

- 基于稀疏线性回归,实现低成本的大规模电路学习
- 在基准数据上达到顶尖干预方法的结构准确性
- 可揭示语义特征在模型中的传播路径,适合模型可解释性研究
机械可解释性领域的一个重要方向是学习大语言模型组件间的稀疏电路,以揭示其协同产生模型行为的机制。然而,原始神经元具有多义性,导致学习到的电路难以解释。稀疏自编码器(SAE)特征缓解了这一问题,但其高维特性使现有基于干预的电路学习方法计算成本过高。我们提出CircuitLasso,一种基于稀疏线性回归的可扩展电路学习方法。该方法在基准数据上实现了与最先进的干预方法相当的结构准确率,同时计算成本大幅降低。在可解释性方面,CircuitLasso能高效发现SAE特征间的关系,揭示人类可理解的语义特征如何在模型中传递并影响预测。最后,我们通过利用所学电路的洞察,在一个领域泛化任务中实现了接近基线性能,且消耗资源显著更低。
原文摘要 · Abstract (English)
A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior. However, raw neurons are polysemantic, making learned circuits hard to interpret. Sparse autoencoder (SAE) features alleviate this, but their high dimensionality makes existing intervention-based circuit learning methods computationally prohibitive. We propose CircuitLasso, a scalable circuit-learning approach based on sparse linear regression. CircuitLasso recovers circuits whose structural accuracy matches that of state-of-the-art intervention-based methods on the benchmark data, at a fraction of the computational cost. For interpretability, CircuitLasso efficiently uncovers relationships among SAE features, showing how human-interpretable semantic features propagate through the model and influence its predictions. Finally, we validate the utility of our learned circuits by leveraging their insights to achieve comparable performance at substantially lower cost on a domain-generalization task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。