arXiv:2606.22994cs.LG2026-06

研究稀疏自编码器能否学出有意义的概念层次结构

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

论文配图:Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
图 1 · 摘自论文原文
  • 提出一套评估概念层次合理性的量化标准
  • 发现特征吸收现象严重破坏层次结构质量
  • 适合关注模型可解释性与特征组织的研究者

稀疏自编码器(SAEs)已成为大型模型中无监督概念发现的重要工具。为提升特征空间的可解释性与可管理性,近期方法开始显式或隐式地引入层次结构,但缺乏统一的评估标准。本文基于语义网络与分类学研究,结合最新SAE工作,提出一组关于泛化/特化层次结构的关键要求,并据此设计具体评估协议。将该协议应用于当前在视觉数据上训练的SAE方法,发现尽管特征空间总体具备构建合理层次的基础,但建立良好层次结构仍具挑战性。特别是特征吸收现象——无论是已知的硬吸收形式,还是连续的软吸收形式——系统性地损害层次质量,揭示了未来方法需面对的根本矛盾。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy, and evaluation has largely relied on qualitative illustrations with fragmented quantitative protocols. To address this, we derive a set of key requirements for generalization/specialization hierarchies in unsupervised concept discovery, drawing on semantic net and taxonomy research alongside recent SAE work, and use them to derive a concrete evaluation protocol. Applying this protocol to current SAE approaches trained on visual data, we find that while feature spaces generally provide a basis for sensible hierarchies, establishing good hierarchical structure remains challenging. In particular, feature absorption, both in its well-known hard form and in a continuous, soft form, systematically compromises hierarchy quality, pointing to a fundamental tension that future approaches will need to navigate.

自编码器概念发现可解释性层次结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。