揭示大模型中特征如何以流形形式存在,解释表示空间距离与概念相关性的关系。
The Origins of Representation Manifolds in Large Language Models
- 将特征视为流形,用表示空间中的最短路径编码特征内在几何结构。
- 在语言模型的嵌入和激活上验证了余弦相似性反映特征连续变化的预测。
- 为理解模型内部表示提供了新视角,适合对可解释性感兴趣的学者。
当前机制可解释性研究致力于将人工智能系统的嵌入和内部表示映射到人类可理解的概念。其中关键假设是线性表征假说,认为神经表示是近乎正交方向向量的稀疏线性组合,反映不同特征的存在与否。该模型支撑了使用稀疏自编码器从表示中恢复特征的方法。近年来,人们开始探讨更完整的特征模型,即神经表示不仅能编码特征是否存在,还能编码其连续且多维的取值。本文阐明了特征为何及如何以流形形式被表示,特别指出表示空间中的余弦相似性可能通过流形上的最短路径编码特征的内在几何结构,从而回答了表示空间距离与概念空间相关性之间的联系问题。理论的关键假设与预测在大型语言模型的文本嵌入和标记激活上得到验证。
原文摘要 · Abstract (English)
There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key element of this effort is the linear representation hypothesis, which posits that neural representations are sparse linear combinations of `almost-orthogonal' direction vectors, reflecting the presence or absence of different features. This model underpins the use of sparse autoencoders to recover features from representations. Moving towards a fuller model of features, in which neural representations could encode not just the presence but also a potentially continuous and multidimensional value for a feature, has been a subject of intense recent discourse. We describe why and how a feature might be represented as a manifold, demonstrating in particular that cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths, potentially answering the question of how distance in representation space and relatedness in concept space could be connected. The critical assumptions and predictions of the theory are validated on text embeddings and token activations of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。