提出可验证的稀疏自编码器解释可信度评估方法
From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

- 用预训练SAE重构激活值构建代理模型,后验评估解释可信性
- 在GPT-2 Small等模型上实现非平凡的可信性边界,后期层更易验证
- 揭示语义对齐与统计稀疏的差异,提供解释可靠性诊断工具
稀疏自编码器(SAE)被广泛用于提取语言模型(LM)中的可解释特征,但核心问题仍是:何时可将SAE生成的解释视为对底层冻结模型的忠实表征?本文通过后验泛化框架,利用替换原生隐藏激活为预训练SAE重建结果所得的稀疏代理模型,对基模型进行可信性认证。该框架基于四个可观测量——代理风险、SAE重建误差、概念池不匹配度和稀疏复杂度——推导出基模型期望风险的上界。该上界可作为解释可信性的操作性标准:非平凡上界表明提取的稀疏特征仍保留有意义的预测信息,而小的重建与不匹配误差则说明代理行为与原模型接近。实证显示,在实际样本量下,GPT-2 Small、Gemma-2B和Llama-3-8B均达到非平凡边界。对Llama-3-8B的逐层分析揭示显著深度依赖性,后期层更易认证,与其更强局部保真性和更弱下游误差放大相关。特征随机置换消融实验表明,该分解能区分真正的语义对齐与仅由统计稀疏导致的假象,为SAE解释的可靠性提供有效诊断。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction. Our framework derives an upper bound on the base model's expected risk using four measurable quantities: proxy risk, SAE reconstruction gap, concept-pool mismatch, and sparse complexity. We interpret this certificate as an operational criterion for explanatory faithfulness. In particular, a non-vacuous bound indicates that the extracted sparse features retain meaningful predictive information, while small reconstruction and mismatch errors indicate that the proxy remains behaviorally close to the original model. Empirically, we show that the bound becomes non-vacuous on GPT-2 Small, Gemma-2B, and Llama-3-8B at practical sample sizes. A detailed layerwise analysis of Llama-3-8B reveals a strong depth dependence, with later layers becoming much easier to certify, associated with both stronger local fidelity and weaker downstream error amplification. Finally, through feature-shuffling ablations, we show that the decomposition distinguishes genuine semantic alignment from mere statistical sparsity, providing a useful diagnostic for when SAE-based explanations become less reliable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。