arXiv:2605.07922cs.LG2026-05

提出树状稀疏自编码器,更准确地学习特征层级结构。

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

论文配图:Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders
图 1 · 摘自论文原文
  • 结合激活与重构双重约束,构建特征层级关系
  • 在多个基准上显著优于现有方法,保持顶尖性能
  • 可揭示大模型中复杂的层级概念结构,适合研究者使用

在稀疏自编码器(SAEs)中学习层级特征对捕捉真实世界数据的结构性至关重要,并能缓解特征吸收或分裂等问题。现有方法依赖激活覆盖率来识别层级关系,假设子特征仅在父特征激活时才触发。然而我们证明该条件不足,常产生语义无关的假阳性。为此,我们引入一种新的重构约束,强化层级间的深层功能关联。通过融合激活与重构双重约束,提出树状稀疏自编码器(Tree SAE),直接从特征集中学习层级结构。实验表明,Tree SAE 在学习层级配对方面显著优于现有 SAE,同时在多个关键基准上保持与最先进方法相当的性能。最后,我们展示了 Tree SAE 在映射子特征子空间几何结构及揭示大语言模型中复杂层级概念结构方面的实际价值。

原文摘要 · Abstract (English)

Learning hierarchical features in Sparse Autoencoders (SAEs) is essential for capturing the structured nature of real-world data and mitigating issues like feature absorption or splitting. Existing works attempt to identify hierarchical relationships within independent feature sets by relying on activation coverage, the assumption that child feature should only activate when its parent feature activates. However, we demonstrate that this condition alone is insufficient; that is, it often produces false positives where parent and child concepts are semantically unrelated. To address this, we introduce a novel reconstruction condition that enforces a deeper functional link between hierarchical levels. By combining both activation and reconstruction constraints, we propose the Tree SAE, a model designed to learn hierarchical structures directly from within the feature set. Our results demonstrate that Tree SAEs significantly surpass the existing SAEs at learning hierarchical pairs while maintaining competitive performance to the state-of-the-art on several key benchmarks. Finally, we demonstrate the practical utility of our Tree SAE in mapping the geometry of child feature subspaces and uncovering the complex hierarchical concept structures encoded within large language models.

稀疏自编码器层级结构特征学习大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。