arXiv:2607.24425cs.LG2026-07

上下文决定模型如何组织概念,能实时塑造几何结构。

Context Is King: How In-Context Specification Shapes the Geometry of Concepts

论文配图:Context Is King: How In-Context Specification Shapes the Geometry of Concepts
图 1 · 摘自论文原文
  • 通过上下文指定规则,模型可动态构建循环或树状结构
  • 大模型中上下文主导的几何结构与激活相似度达0.6–0.9
  • 小模型虽有粗略结构但无法干净使用,规模关键

大型语言模型将结构化概念置于几何上忠实的流形上:工作日形成圆环,月份构成另一圆环,通常视为网络存储并查表的固定世界模型。本文表明:上下文为王——模型实际使用的结构由上下文指定决定。一个陈述性规则不仅确定几何编码的关系,还决定其拓扑类型:相同词元可按指令形成循环或分支树,甚至在无意义、无先验的词元上构建,这是重新标记存储形状所无法实现的。当指定结构与强预训练先验冲突时,上下文设定的几何结构在能力强的模型中占据主导(表示相似度0.6–0.9,远高于先验近零),跨测试先验和两个模型家族(Gemma、Qwen)均成立。激活修补显示该映射是因果使用的,而非探测相关:交换某实体的激活使模型以另一实体的后继作答。粗略地图在小模型中即存在,但能否干净使用取决于规模:清晰主导与因果切换仅在更大模型(至Gemma-31B、Qwen-27B)中出现,小于此规模则弱化或反转,因此同家族的大模型可能具备而小模型缺失该机制。模型是全新构建还是重构存储结构尚不明确;操作上,模型所用几何即上下文指定者。

原文摘要 · Abstract (English)

Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed world-model the network stores and looks up. We show that context is king: the structure a model actually uses is set by the in-context specification. A declarative rule fixes not only which relations the geometry encodes but its topology type: the same tokens form a cycle or a branching tree on command, built even on arbitrary, meaning-free tokens with no prior to inherit, which a relabeled stored shape cannot do. When the specification conflicts with a strong pretrained prior, the context-set geometry dominates it in capable models, read from the same activations (representational similarity 0.6--0.9 to the imposed structure versus near-zero to the prior), across the priors we test and both families we study (Gemma, Qwen). Activation patching shows the map is causally used, not a probe correlate: swapping one entity's activation for another's makes the model answer with the other entity's successor under the imposed order. A rough map forms readily, present even in small and base models; what scale gates is using it cleanly: clean dominance and the causal crossover emerge only in the larger models (up to Gemma-31B and Qwen-27B) and weaken or reverse below, so a mechanism present in a large model can be absent in a smaller one of the same family. Whether the model builds this geometry anew or reconfigures a stored one we leave open; operationally, the geometry it uses is the one the context specifies.

上下文建模几何结构大模型机制激活分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。