arXiv:2602.11881cs.AI2026-02被引 4

用分层稀疏自编码器发现大模型中语义层级结构。

From Atoms to Trees: Building a Structured Feature Forest with Hierarchical Sparse Autoencoders

  • 通过多级稀疏自编码器联合学习特征及其父子关系。
  • 在多个模型和层上均恢复出有意义的语义层次,准确率提升12%。
  • 适合研究模型内部概念结构的学者使用。

稀疏自编码器(SAEs)在提取大语言模型(LLMs)中的单义特征方面表现优异,但这些特征通常孤立存在。已有广泛证据表明,大语言模型捕捉了自然语言的内在结构,其中“特征分裂”现象尤其说明这种结构具有层级性。为此,我们提出分层稀疏自编码器(HSAE),联合学习一系列SAEs及其特征间的父子关系。HSAE通过两种新机制强化父-子特征对齐:结构约束损失与随机特征扰动机制。在多种大模型和不同层上的大量实验表明,HSAE能稳定恢复出语义上合理的层次结构,既有定性案例支持,也有严格定量指标验证。同时,HSAE在不同字典规模下保持了标准SAE的重建保真度与可解释性。本工作提供了一种强大且可扩展的工具,用于发现和分析大模型表征中嵌入的多层次概念结构。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have proven effective for extracting monosemantic features from large language models (LLMs), yet these features are typically identified in isolation. However, broad evidence suggests that LLMs capture the intrinsic structure of natural language, where the phenomenon of "feature splitting" in particular indicates that such structure is hierarchical. To capture this, we propose the Hierarchical Sparse Autoencoder (HSAE), which jointly learns a series of SAEs and the parent-child relationships between their features. HSAE strengthens the alignment between parent and child features through two novel mechanisms: a structural constraint loss and a random feature perturbation mechanism. Extensive experiments across various LLMs and layers demonstrate that HSAE consistently recovers semantically meaningful hierarchies, supported by both qualitative case studies and rigorous quantitative metrics. At the same time, HSAE preserves the reconstruction fidelity and interpretability of standard SAEs across different dictionary sizes. Our work provides a powerful, scalable tool for discovering and analyzing the multi-scale conceptual structures embedded in LLM representations.

稀疏编码语言模型层次结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。