arXiv:2507.08473cs.LG2025-07被引 3

不依赖自然语言解释,直接评估稀疏自编码器的可解释性。

Evaluating SAE interpretability without explanations

  • 绕过生成解释的步骤,直接衡量潜在表示的可解释性。
  • 与人类评估结果一致,验证了新方法的有效性。
  • 为可解释性评估提供更标准、更可信的基准方法。

稀疏自编码器(SAEs)和转换器已成为机器学习可解释性的重要工具。然而,衡量其可解释性仍具挑战,且缺乏公认的评估基准。现有方法通常先为每个潜在变量生成一句自然语言解释,再通过大模型预测其在新上下文中的激活情况来评估解释质量。这种方法难以将解释生成过程与潜在变量本身的实际可解释性区分开。本文改进现有方法,提出一种无需生成自然语言解释即可评估稀疏编码器可解释性的新框架,实现更直接、更标准化的评估。此外,我们将所提指标得分与人类评估结果进行对比,涵盖相似任务和不同设置,为社区优化此类技术的评估提供参考建议。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) and transcoders have become important tools for machine learning interpretability. However, measuring how interpretable they are remains challenging, with weak consensus about which benchmarks to use. Most evaluation procedures start by producing a single-sentence explanation for each latent. These explanations are then evaluated based on how well they enable an LLM to predict the activation of a latent in new contexts. This method makes it difficult to disentangle the explanation generation and evaluation process from the actual interpretability of the latents discovered. In this work, we adapt existing methods to assess the interpretability of sparse coders, with the advantage that they do not require generating natural language explanations as an intermediate step. This enables a more direct and potentially standardized assessment of interpretability. Furthermore, we compare the scores produced by our interpretability metrics with human evaluations across similar tasks and varying setups, offering suggestions for the community on improving the evaluation of these techniques.

可解释性稀疏编码评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。