揭示大模型在上下文学习中同时依赖结构推理与局部模式匹配。
Belief or Circuitry? Causal Evidence for In-Context Graph Learning

- 通过主成分分析发现模型同时编码两种图结构的全局拓扑信息。
- 晚期层插补实验显示可几乎完全转移对正确图结构的偏好。
- 适用于研究大模型内部机制、因果推理与上下文学习的读者。
大模型如何在上下文学习中进行推理?是仅模仿最近的词元模式,还是推断潜在结构?我们通过一个跨两种竞争图结构的简化随机游走任务探究此问题。该任务的答案原则上可区分:模型要么追踪全局拓扑,要么复制局部转移。我们提出两条证据表明单一机制不足以解释现象。首先,通过主成分分析(PCA)重构内部表征结构发现,在中间混合比例下,两种图结构同时以正交主子空间编码。这一模式难以用纯局部转移复制解释。其次,残差流激活插补与图差异引导的因果干预表明:晚期层插补几乎完全转移了对干净图结构的偏好;线性引导能按预期方向改变预测,但在归一化匹配和标签打乱控制下失效。综合来看,结果最支持双机制模型:真正的结构推断与归纳电路并行运作。
原文摘要 · Abstract (English)
How do LLMs learn in-context? Is it by pattern-matching recent tokens, or by inferring latent structure? We probe this question using a toy graph random-walk across two competing graph structures. This task's answer is, in principle, decidable: either the model tracks global topology, or it copies local transitions. We present two lines of evidence that neither account alone is sufficient. First, reconstructing the internal representation structure via PCA reveals that at intermediate mixture ratios, both graph topologies are encoded in orthogonal principal subspaces simultaneously. This pattern is difficult to reconcile with purely local transition copying. Second, residual-stream activation patching and graph-difference steering causally intervene on this graph-family signal: late-layer patching almost fully transfers the clean graph preference, while linear steering moves predictions in the intended direction and fails under norm-matched and label-shuffled controls. Taken together, our findings are most consistent with a dual-mechanism account in which genuine structure inference and induction circuits operate in parallel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。