发现大模型用螺旋结构存储化学知识,中间层实现间接回忆。
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
- 通过隐藏状态分析揭示周期表的三维螺旋几何结构。
- 中层编码连续重叠属性,支持间接记忆;深层强化分类与语境。
- 为科学知识建模提供新视角,适合材料科学等领域研究者。
本研究以化学元素和LLaMA系列模型为例,探索大语言模型如何编码交织的科学知识。我们发现在隐藏状态中存在与周期表概念结构对齐的3D螺旋结构,表明模型能从文本中学习到科学概念的几何组织。线性探测显示,中间层编码连续且重叠的属性,支持间接回忆;深层则强化类别区分并融合语言上下文。结果表明,大语言模型并非孤立存储事实,而是将符号知识表示为跨层交织的语义流形。该研究旨在启发对大模型如何表征与推理科学知识的进一步探索,尤其在材料科学等领域的应用。
原文摘要 · Abstract (English)
This study explores how large language models (LLMs) encode interwoven scientific knowledge, using chemical elements and LLaMA-series models as a case study. We identify a 3D spiral structure in the hidden states that aligns with the conceptual structure of the periodic table, suggesting that LLMs can reflect the geometric organization of scientific concepts learned from text. Linear probing reveals that middle layers encode continuous, overlapping attributes that enable indirect recall, while deeper layers sharpen categorical distinctions and incorporate linguistic context. These findings suggest that LLMs represent symbolic knowledge not as isolated facts, but as structured geometric manifolds that intertwine semantic information across layers. We hope this work inspires further exploration of how LLMs represent and reason about scientific knowledge, particularly in domains such as materials science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。