arXiv:2602.02991cs.AI2026-02

LLM推理时看似短视,实则因上下文积累导致规划行为变化。

Large Language Models Can Take False First Steps at Inference-time Planning

  • 用贝叶斯框架解释上下文演化如何影响LLM的规划行为
  • 自生成内容积累后,规划能力显著增强,初始偏差降低
  • 适合关注LLM推理机制与规划缺陷的研究者

大型语言模型(LLMs)在训练中习得序列级规划能力,但推理时的规划行为常显得短视且不一致。本文提出一种贝叶斯解释:规划行为根植于不断演化的生成上下文。由于自然语言与模型内化语言的细微差异,自生成上下文的积累会引发推理过程中的规划转移,从而造成规划能力减弱的假象。通过两项受控实验验证:在随机生成任务中,人类提示下规划受限,但随着自生成上下文累积,规划强度提升;在高斯采样任务中,基于自生成序列的条件化可减少初始偏差。这些发现为理解LLM推理时的前瞻性规划提供了理论解释与实证支持。

原文摘要 · Abstract (English)

Large language models (LLMs) have been shown to acquire sequence-level planning abilities during training, yet their planning behavior exhibited at inference time often appears short-sighted and inconsistent with these capabilities. We propose a Bayesian account for this gap by grounding planning behavior in the evolving generative context: given the subtle differences between natural language and the language internalized by LLMs, accumulated self-generated context drives a planning-shift during inference and thereby creates the appearance of compromised planning behavior. We further validate the proposed model through two controlled experiments: a random-generation task demonstrating constrained planning under human prompts and increasing planning strength as self-generated context accumulates, and a Gaussian-sampling task showing reduced initial bias when conditioning on self-generated sequences. These findings provide a theoretical explanation along with empirical evidence for characterizing how LLMs plan ahead during inference.

大模型推理规划能力上下文演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。