arXiv:2506.01197cs.CLcs.AI2025-06被引 11

让稀疏自编码器学会概念间的语义层次关系,提升解释性和效率。

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

  • 在稀疏自编码器中显式建模概念的语义层级结构。
  • 在大语言模型表征上实现更优的重建效果与可解释性。
  • 显著提升计算效率,适合需要高效解释的AI系统。

稀疏字典学习(尤其是稀疏自编码器)旨在学习一组可理解的概念,以解释抽象空间中的变化。其基本局限在于无法利用或表达所学概念之间的语义关系。本文提出一种改进的SAE架构,显式建模概念的语义层次结构。应用于大型语言模型的内部表示时,结果表明语义层次能够被有效学习,且能同时提升重建质量与可解释性。此外,该架构显著提升了计算效率。

原文摘要 · Abstract (English)

Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.

稀疏自编码器语义层次可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。