arXiv:2602.01687cs.CLcs.AI2026-02

发现大模型用子空间向量运算解决任务,揭示其隐式推理机制。

Functional Subspace, where language models can use vector algebra to solve problems

  • 通过分析上下文学习中的激活流,发现模型构建可积累证据的子空间。
  • 在子空间内仅用简单向量加减即可完成复杂任务求解。
  • 为理解模型隐式推理提供新视角,适合研究模型可解释性者阅读。

大语言模型(LLMs)最初用于自然语言任务,但已展现出跨领域执行复杂功能的能力,并能在未显式训练的情况下习得新技能。因此,深入理解其运行机制与局限对诊断和修复至关重要。已有研究提出,高级概念在模型激活空间中以线性方向编码,嵌入空间的几何结构具有语义意义。受此启发,我们假设LLMs可能利用子空间和子空间内的向量代数来执行任务。为此,我们分析了在上下文学习(ICL)过程中,模型的功能模块与残差流。结果表明:1)模型能够创建子空间,用于累积证据;2)通过子空间中的简单代数操作即可解决ICL任务。这揭示了模型潜在的隐式计算机制。

原文摘要 · Abstract (English)

Large language models (LLMs) were invented for natural language tasks such as translation, but they have proved that they can perform highly complex functions across domains. Additionally, they have been thought to develop new skills without being trained on them. These learning capabilities lead to LLMs adoption in a wide range of domains. Thus, it is imperative that we understand their operating mechanisms and limitations for proper diagnostics and repair. The earlier studies proposed that high level concepts are encoded as linear directions in LLMs activation space and that the geometry of embeddings have semantic meanings. Inspired by these studies, we hypothesize that LLMs may use subspaces and vector algebra in subspaces to perform tasks. To address this hypothesis, we analyze LLMs' functional modules and residual streams collected from LLMs engaging in in-context learning (ICL), one of the emergent abilities. Our analyses suggest that 1) LLMs can create subspaces, where evidence can be accumulated and 2) ICL tasks can be solved via simple algebraic operations in subspaces.

大模型子空间向量代数可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。