发现大模型内部存在类海马体的抽象几何表征,支持推理能力。
Abstract representational geometry supports inference in large language models
- 通过文本化反转学习任务,对比人类与大模型的表征结构。
- 高层神经元形成类海马体的抽象几何结构,与推理相关。
- 几何结构可调控:正则化提升泛化推理,语言建模促解耦。
人类智能的核心在于从稀疏观察中推断潜在任务结构以适应变化环境。神经科学表明,这一能力依赖于海马体构建抽象表征,表现为神经状态空间中的低维、近正交流形。然而,大语言模型(LLMs)的内部机制仍不透明,尚不清楚其是否具备类似抽象表征,还是仅依赖特定任务的统计规律进行推理。本文将情境反转学习范式转化为文本任务,从行为与表征层面比较人类与LLMs。结果表明,尽管大模型泛化推理能力低于人类,但当推理发生时,其内部状态展现出与海马体报告相似的抽象几何结构。值得注意的是,这种表征几何并非均匀分布,而是沿模型深度呈层次组织:低层稳定编码刺激身份,高层则形成富含抽象上下文几何的类海马功能带。此外,干预实验进一步揭示几何结构在推理中的作用:任务序列语言建模诱导几何解耦,而对高层施加几何正则化可显著提升泛化推理的出现频率。这些发现确立了抽象表征几何作为大语言模型推理的机制性原理。
原文摘要 · Abstract (English)
A defining feature of human intelligence is the ability to adapt to changing environments by inferring latent task structure from sparse observations. Neuroscientific research indicates that this capability relies on the hippocampus constructing abstract representations, expressed as low-dimensional, approximately orthogonal manifolds in neural state space. However, the internal mechanisms of large language models (LLMs) remain largely opaque, making it unclear whether they form comparable abstract representations or instead rely on task-specific statistical regularities when performing comparable reasoning tasks. Here we adapt a contextual reversal-learning paradigm to a text-based setting and compare humans and LLMs at both the Behavioural and representational levels. We report that although LLMs exhibit generalizable reasoning less frequently than humans, when such inference occurs, their internal states exhibit abstract geometric structures that resemble those reported in the hippocampus. Notably, this representational geometry is not uniformly distributed but is organized hierarchically across model depth: whereas lower layers show early, stable encoding of stimulus identity, higher layers form a hippocampal-like functional band enriched for abstract context geometry associated with inference. Furthermore, complementary intervention experiments mechanistically implicate geometry in reasoning: task-sequence language modelling induces geometric disentanglement, whereas geometric regularization of higher layers increases the emergence of generalizable inference. Together, these findings establish abstract representational geometry as a mechanistic principle supporting inference in large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。