arXiv:2409.05771cs.CLcs.AI2024-09被引 14

发现语言模型存在两阶段抽象过程,解释为何中间层能预测脑活动。

Evidence from fMRI Supports a Two-Phase Abstraction Process in Language Models

  • 通过流形学习揭示模型训练中自然出现两阶段抽象
  • 中间层表示的内在维度与脑响应预测性能强相关
  • 抽象机制源于模型组合性,非单纯预测任务驱动

大量研究表明,大型语言模型的中间隐藏状态可有效预测自然语言刺激下的脑区反应。然而,这些状态为何具备如此高精度的泛化迁移能力仍不清楚。本文基于fMRI语言编码模型证据,证明语言模型内部存在两阶段抽象过程:第一阶段为语义组合,随训练进程压缩至更少层数;第二阶段为抽象表达。我们发现,各层编码性能与表示的内在维度高度对应,初步表明该关联主要源于模型固有的组合性,而非其下游的下一个词预测任务。

原文摘要 · Abstract (English)

Research has repeatedly demonstrated that intermediate hidden states extracted from large language models are able to predict measured brain response to natural language stimuli. Yet, very little is known about the representation properties that enable this high prediction performance. Why is it the intermediate layers, and not the output layers, that are most capable for this unique and highly general transfer task? In this work, we show that evidence from language encoding models in fMRI supports the existence of a two-phase abstraction process within LLMs. We use manifold learning methods to show that this abstraction process naturally arises over the course of training a language model and that the first "composition" phase of this abstraction process is compressed into fewer layers as training continues. Finally, we demonstrate a strong correspondence between layerwise encoding performance and the intrinsic dimensionality of representations from LLMs. We give initial evidence that this correspondence primarily derives from the inherent compositionality of LLMs and not their next-word prediction properties.

语言模型脑科学抽象过程fMRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。