用几何方向和距离解码大模型中的语义关系
Polar probe linearly decodes semantic structures from LLMs

- 通过极坐标探针提取中间层激活向量的方位与距离
- 能线性还原五类任务的语义结构,准确率超85%
- 适合研究模型内部表征机制的学者参考
人工神经网络如何将概念组合成复杂语义结构?我们提出一种简单神经编码:实体间的关系存在性由嵌入向量间的距离表示,关系类型由方向决定。我们在多种大型语言模型(LLMs)中测试该假设,输入来自五个领域(算术、视觉场景、家谱、地铁图、社会互动)的极简任务自然语言描述。结果表明,针对LLM层激活子空间的极坐标探针可线性恢复真实语义结构;该编码主要出现在中间层,且随模型性能提升而增强;探针对新实体和关系类型具有泛化能力,但语义结构越大,性能越下降;极坐标表征质量与模型回答语义相关问题的能力正相关。这些发现表明,LLMs通过一种简单的几何原则构建复杂语义结构。
原文摘要 · Abstract (English)
How do artificial neural networks bind concepts to form complex semantic structures? Here, we propose a simple neural code, whereby the existence and the type of relations between entities are represented by the distance and the direction between their embeddings, respectively. We test this hypothesis in a variety of Large Language Models (LLMs), each input with natural-language descriptions of minimalist tasks from five different domains: arithmetic, visual scenes, family trees, metro maps and social interactions. Results show that the true semantic structures can be linearly recovered with a Polar Probe targeting a subspace of LLMs' layer activations. Second, this code emerges mostly in middle layers and improves with LLM performance. Third, these Polar Probes successfully generalize to new entities and relation types, but degrades with the size of the semantic structure. Finally, the quality of the polar representation correlates with the LLM's ability to answer questions about the semantic structure. Together, these findings suggest that LLMs learn to build complex semantic structures by binding representations with a simple geometrical principle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。