arXiv:2410.19750q-bio.NCcs.AI2024-10被引 66

揭示大模型概念特征的三维几何结构:原子、脑区与星系级规律

The Geometry of Concepts: Sparse Autoencoder Feature Structure

  • 通过稀疏自编码器解析概念向量,发现其呈现晶体状小尺度结构
  • 中层特征在空间上形成数学与代码等模块化聚集区,类似大脑功能区
  • 深层特征呈现幂律分布,中间层特征聚类熵最低,具显著层级结构

稀疏自编码器最近生成了对应大型语言模型概念宇宙的高维向量字典。我们发现该概念宇宙在三个层面具有有趣结构:1)‘原子’小尺度结构包含面为平行四边形或梯形的‘晶体’,广义推广了(男-女-王-后)等经典例子;当通过线性判别分析剔除词长等全局干扰方向后,这些平行四边形及其关联函数向量质量显著提升。2)‘脑’中间尺度结构表现出显著的空间模块性,例如数学与代码特征形成类似神经功能性脑叶的‘叶瓣’;我们用多种指标量化其空间局部性,发现粗粒度下共现特征簇在空间上聚集程度远超随机预期。3)‘星系’大尺度特征点云非各向同性,特征值呈幂律分布,中间层斜率最陡;同时量化了不同层的聚类熵变化。

原文摘要 · Abstract (English)

Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that this concept universe has interesting structure at three levels: 1) The "atomic" small-scale structure contains "crystals" whose faces are parallelograms or trapezoids, generalizing well-known examples such as (man-woman-king-queen). We find that the quality of such parallelograms and associated function vectors improves greatly when projecting out global distractor directions such as word length, which is efficiently done with linear discriminant analysis. 2) The "brain" intermediate-scale structure has significant spatial modularity; for example, math and code features form a "lobe" akin to functional lobes seen in neural fMRI images. We quantify the spatial locality of these lobes with multiple metrics and find that clusters of co-occurring features, at coarse enough scale, also cluster together spatially far more than one would expect if feature geometry were random. 3) The "galaxy" scale large-scale structure of the feature point cloud is not isotropic, but instead has a power law of eigenvalues with steepest slope in middle layers. We also quantify how the clustering entropy depends on the layer.

稀疏自编码器概念几何特征结构语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。