用聚焦概念的可视化方法,更清晰地探索大模型中稀疏自编码器的特征关系。
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
- 聚焦人工标注的概念,只展示相关特征,避免全量特征可视化混乱。
- 结合拓扑编码与降维技术,真实呈现局部与全局特征关系。
- 适合研究大模型内部表征机制的科研人员使用。
稀疏自编码器(SAEs)已成为揭示大型语言模型(LLMs)中可解释特征的重要工具,通过学习稀疏方向实现。然而,提取出的方向数量庞大,全面探索难以实现。传统嵌入方法如UMAP虽能揭示全局结构,但存在高维压缩失真、点重叠和邻域误判等局限。本文提出一种聚焦式探索框架,优先展示人工标注的概念及其对应的SAE特征,而非试图同时可视化所有特征。我们构建了一个交互式可视化系统,融合基于拓扑的视觉编码与降维技术,准确呈现选定特征间的局部与全局关系。该混合方法使用户可通过有目标的可解释子集深入分析隐空间中的概念表征行为。
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) have emerged as a powerful tool for uncovering interpretable features in large language models (LLMs) through the sparse directions they learn. However, the sheer number of extracted directions makes comprehensive exploration intractable. While conventional embedding techniques such as UMAP can reveal global structure, they suffer from limitations including high-dimensional compression artifacts, overplotting, and misleading neighborhood distortions. In this work, we propose a focused exploration framework that prioritizes curated concepts and their corresponding SAE features over attempts to visualize all available features simultaneously. We present an interactive visualization system that combines topology-based visual encoding with dimensionality reduction to faithfully represent both local and global relationships among selected features. This hybrid approach enables users to investigate SAE behavior through targeted, interpretable subsets, facilitating deeper and more nuanced analysis of concept representation in latent space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。