arXiv:2605.12412cs.CLcs.AI2026-05被引 1

用几何空间解释大模型如何动态更新信念。

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space

论文配图:Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
图 1 · 摘自论文原文
  • 将模型信念变化视为低维流形上的轨迹,揭示内在结构。
  • 行为与内部表征均体现一致的几何模式,可用简单探测器预测。
  • 操纵表征可定向引导信念演变,效果由空间几何决定。

大型语言模型(LLMs)在上下文中会动态调整其行为,可视为贝叶斯推理过程。然而,该推理所依赖的潜在假设空间结构仍不清晰。本文提出,LLMs 在一个低维几何空间——概念信念空间中分配信念,且上下文学习对应于该空间中的信念轨迹演化。以故事理解为动态信念更新的自然场景,结合行为分析与表征分析研究这些轨迹。结果发现:(1) 信念更新可被良好描述为低维、有结构的流形上的轨迹;(2) 该结构在模型行为与内部表征中均一致呈现,且可通过简单的线性探测器预测行为;(3) 对表征的干预能因果性地引导信念轨迹,其影响可由概念空间的几何结构预测。综合来看,本研究为LLM中的信念动态提供了几何解释,将贝叶斯视角的上下文学习建立在结构化的概念表征基础上。

原文摘要 · Abstract (English)

Large Language Models (LLMs) update their behavior in context, which can be viewed as a form of Bayesian inference. However, the structure of the latent hypothesis space over which this inference operates remains unclear. In this work, we propose that LLMs assign beliefs over a low-dimensional geometric space - a conceptual belief space - and that in-context learning corresponds to a trajectory through this space as beliefs are updated over time. Using story understanding as a natural setting for dynamic belief updating, we combine behavioral and representational analyses to study these trajectories. We find that (1) belief updates are well-described as trajectories on low-dimensional, structured manifolds; (2) this structure is reflected consistently in both model behavior and internal representations and can be decoded with simple linear probes to predict behavior; and (3) interventions on these representations causally steer belief trajectories, with effects that can be predicted from the geometry of the conceptual space. Together, our results provide a geometric account of belief dynamics in LLMs, grounding Bayesian interpretations of in-context learning in structured conceptual representations.

大模型信念空间几何表征上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。