用层次嵌入可视化语言模型的稀疏特征,方便发现概念间的关联。
Navigating the Concept Space of Language Models
- 构建多分辨率流形,通过层级邻域嵌入组织稀疏特征
- 在SmolLM2上发现高阶结构、有意义子簇和罕见概念
- 适合研究模型内部表示的学者快速探索特征语义
在大规模语言模型激活上训练的稀疏自编码器(SAEs)会产生数千个可映射到人类可理解概念的特征。当前分析这些特征主要依赖查看激活最强的例子、手动浏览单个特征或对感兴趣概念进行语义搜索,难以规模化地进行探索性发现。本文提出Concept Explorer,一个可扩展的交互式后处理系统,利用层级邻域嵌入组织概念解释。该方法在SAE特征嵌入上构建多分辨率流形,支持从粗粒度概念簇逐步导航到细粒度邻域,实现概念的发现、比较与关系分析。我们在SmolLM2提取的SAE特征上验证了其有效性,揭示了连贯的高层结构、有意义的子簇以及现有流程难以识别的独特罕见概念。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) trained on large language model activations output thousands of features that enable mapping to human-interpretable concepts. The current practice for analyzing these features primarily relies on inspecting top-activating examples, manually browsing individual features, or performing semantic search on interested concepts, which makes exploratory discovery of concepts difficult at scale. In this paper, we present Concept Explorer, a scalable interactive system for post-hoc exploration of SAE features that organizes concept explanations using hierarchical neighborhood embeddings. Our approach constructs a multi-resolution manifold over SAE feature embeddings and enables progressive navigation from coarse concept clusters to fine-grained neighborhoods, supporting discovery, comparison, and relationship analysis among concepts. We demonstrate the utility of Concept Explorer on SAE features extracted from SmolLM2, where it reveals coherent high-level structure, meaningful subclusters, and distinctive rare concepts that are hard to identify with existing workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。