用类别向量空间解释神经元多义性,提升语言模型效率
Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces
- 将神经元定义为非正交基的类别向量空间,基于前层子维度构建
- 通过神经内注意力识别关键类别区域,提升模型运行效率
- 为理解神经元多义性提供新几何视角,适合模型可解释性研究者
当前对语言模型中合成神经元多义性的理解,认为是潜在空间中分布式特征叠加的必然结果。本文提出一种新方法:将第n层神经元几何定义为由第n-1层神经元提取的类别子维度组成的非正交基类别向量空间。该空间由每个神经元的激活空间结构化,并通过神经内注意力过程,识别并利用一个关键类别区域——该区域更均匀且位于多个类别子维度的交集中,从而提升语言模型的效率。
原文摘要 · Abstract (English)
The polysemantic nature of synthetic neurons in artificial intelligence language models is currently understood as the result of a necessary superposition of distributed features within the latent space. We propose an alternative approach, geometrically defining a neuron in layer n as a categorical vector space with a non-orthogonal basis, composed of categorical sub-dimensions extracted from preceding neurons in layer n-1. This categorical vector space is structured by the activation space of each neuron and enables, via an intra-neuronal attention process, the identification and utilization of a critical categorical zone for the efficiency of the language model - more homogeneous and located at the intersection of these different categorical sub-dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。