arXiv:2605.18830cs.LG2026-05被引 3

揭示了上下文学习中任务信息集中在低维概念子空间内。

In-Context Learning Operates as Concept Subspace Learning

  • 将上下文学习视为在概念子空间中的推理,任务变化仅沿内在概念坐标。
  • 68–73维子空间可恢复78.8%的准确率差距,互补子空间无效。
  • 适合研究大模型内部机制与可控推理的学者参考。

本文从概念子空间视角研究上下文学习(ICL),认为任务变化仅沿内在概念坐标进行,尽管输入处于高维空间。对岭回归和最小二乘法的ICL代理模型分析表明,预测可精确分解为概念坐标回归与非子空间泄漏项。在块对角或近似块对角协方差假设下,主要估计误差与干扰敏感性与概念子空间维度相关,残差效应由跨子空间耦合控制。实验发现,在基于CounterFact的多关系提示中,使用Llama-3-8B模型时,4096维残差流中的68–73维子空间能恢复78.8%的干净-污染准确率差距,而补全子空间无贡献。概念替换使预测转向注入关系,随机及跨任务匹配秩对照组效果甚微。在Qwen2.5-7B和受控跨语言规则任务中亦呈现相同模式。结果支持概念子空间作为结构化任务族中可恢复的、任务对齐的中介,但不意味着完整电路恢复。

原文摘要 · Abstract (English)

Regression and Bayesian accounts of in-context learning (ICL) explain how demonstrations can induce predictors, while mechanistic analyses often identify compact activation directions that steer prompted behavior. However, it remains unclear whether structured demonstrations induce low-dimensional concept inference. We study this question through a concept-subspace view of ICL, in which tasks vary only along intrinsic concept coordinates, although inputs are observed in a high-dimensional ambient space. For ridge and least-squares ICL proxies, prediction decomposes exactly into concept-coordinate regression and off-subspace leakage. Under block-diagonal or near-block-diagonal covariance assumptions, the leading estimation and nuisance-sensitivity terms scale with the dimension of the concept subspace, while residual effects are controlled by cross-subspace coupling. This separation gives a mechanistic prediction: recoverable task information should concentrate in a low-dimensional, task-aligned activation subspace. On CounterFact-derived multi-relation prompts with Llama-3-8B, a 68--73-dimensional subspace of the 4096-dimensional residual stream restores 78.8% of the clean--corrupted accuracy gap, whereas patching the complementary subspace restores 0%. Concept swaps redirect predictions toward injected relations, while random and cross-task matched-rank controls are largely ineffective. Additional experiments on Qwen2.5-7B and a controlled cross-lingual rule task show the same qualitative pattern. These results support concept subspaces as compact, task-aligned mediators of recoverable ICL behavior in structured task families, without implying full-circuit recovery.

上下文学习概念子空间大模型机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。