arXiv:2502.11932cs.CL2025-02

发现大模型中语言与算术在表征空间完全分离,类似人脑神经机制。

On Representational Dissociation of Language and Arithmetic in Large Language Models

  • 通过线性分类器和聚类可分性测试分析表征几何结构
  • 算术题与普通语言输入在所有层都位于独立区域
  • 结果暗示算术推理有专属表征区,适合认知建模研究者

人类语言与非语言思维能力之间的关联长期存在争议,近年来神经科学证据提供了支持。这一科学背景自然引出一个跨学科问题:大语言模型中是否存在类似的语言-思维分离?本文首次探索此问题,聚焦简单算术能力(如1+2=?)作为思维能力,分析其在模型内部表征空间中的几何结构。通过线性分类器与聚类可分性测试,发现算术方程与一般语言输入在所有层的表征空间中均位于完全分离区域,该结论在更受控刺激(如拼写出的方程)下仍成立。初步表明算术推理可能被映射到与通用语言输入不同的独立区域,与人脑激活模式一致;但我们也指出其存在某些认知上不合理的几何特性。

原文摘要 · Abstract (English)

The association between language and (non-linguistic) thinking ability in humans has long been debated, and recently, neuroscientific evidence of brain activity patterns has been considered. Such a scientific context naturally raises an interdisciplinary question -- what about such a language-thought dissociation in large language models (LLMs)? In this paper, as an initial foray, we explore this question by focusing on simple arithmetic skills (e.g., $1+2=$ ?) as a thinking ability and analyzing the geometry of their encoding in LLMs' representation space. Our experiments with linear classifiers and cluster separability tests demonstrate that simple arithmetic equations and general language input are encoded in completely separated regions in LLMs' internal representation space across all the layers, which is also supported with more controlled stimuli (e.g., spelled-out equations). These tentatively suggest that arithmetic reasoning is mapped into a distinct region from general language input, which is in line with the neuroscientific observations of human brain activations, while we also point out their somewhat cognitively implausible geometric properties.

大模型表征算术推理认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。