arXiv:2411.00743cs.LGcs.AI2024-11NAACL被引 17

针对大模型中罕见概念难捕捉的问题,提出专用稀疏自编码器提升可解释性。

Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models

  • 为特定子领域设计专用稀疏自编码器,聚焦罕见概念
  • 在偏见生物数据集上提升最差群体准确率12.5%
  • 适用于需要深入理解模型内部机制的研究者

理解并缓解基础模型(FMs)潜在风险依赖有效的可解释性方法。稀疏自编码器(SAEs)虽能解耦模型表征,却难以捕捉数据中罕见但关键的概念。本文提出专用稀疏自编码器(SSAEs),通过聚焦特定子领域来揭示这些“暗物质”特征。我们提供了一套实用的训练方案,证明密集检索在数据选择中的有效性,以及倾斜经验风险最小化作为训练目标可提升概念召回率。在下游困惑度和 $L_0$ 稀疏度等标准指标上,SSAEs 在捕捉子领域尾部概念方面优于通用 SAEs。在偏见生物数据集的案例研究中,应用 SSAEs 去除虚假性别信息后,最差群体分类准确率提升 12.5%。SSAEs 为深入探查基础模型在子领域中的内在运作提供了新视角。

原文摘要 · Abstract (English)

Understanding and mitigating the potential risks associated with foundation models (FMs) hinges on developing effective interpretability methods. Sparse Autoencoders (SAEs) have emerged as a promising tool for disentangling FM representations, but they struggle to capture rare, yet crucial concepts in the data. We introduce Specialized Sparse Autoencoders (SSAEs), designed to illuminate these elusive dark matter features by focusing on specific subdomains. We present a practical recipe for training SSAEs, demonstrating the efficacy of dense retrieval for data selection and the benefits of Tilted Empirical Risk Minimization as a training objective to improve concept recall. Our evaluation of SSAEs on standard metrics, such as downstream perplexity and $L_0$ sparsity, show that they effectively capture subdomain tail concepts, exceeding the capabilities of general-purpose SAEs. We showcase the practical utility of SSAEs in a case study on the Bias in Bios dataset, where SSAEs achieve a 12.5\% increase in worst-group classification accuracy when applied to remove spurious gender information. SSAEs provide a powerful new lens for peering into the inner workings of FMs in subdomains.

可解释性稀疏编码基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。