arXiv:2604.12493cs.CLcs.AI2026-04被引 5

大模型越大会自动规划,且能提前布局未来内容。

Latent Planning Emerges with Scale

  • 定义隐式规划:模型内部预存未来内容线索并引导当前输出。
  • 0.6B到14B的模型显示规模越大,规划能力越强,4B-8B已具雏形。
  • 在押韵诗中模型常提前识韵,但深度规划仍有限,可被引导增强。

大语言模型(LLMs)能完成看似需规划的任务,如撰写连贯故事或生成代码,却未显式表达计划;其隐式规划能力尚不明确。本文将隐式规划定义为:模型内部存在对未来某个词或概念的表示,并影响先前上下文以支持该未来内容的生成。我们研究了Qwen-3系列(0.6B–14B)在简单规划任务上的表现,发现规划能力随模型规模增加而提升。具备规划能力的模型会预先表征如“accountant”这样的目标词,并导致输出“an”而非“a”;即使较小的Qwen-3 4B–8B模型也展现出初步规划机制。在更复杂的押韵对句任务中,模型常能提前识别韵脚,但很少进行远期规划。然而,通过引导模型关注特定计划词,可在散文生成中激发更多规划行为,且该行为随规模增长。本研究提出测量规划的框架,并提供了机制证据,说明模型规划能力随规模演进。

原文摘要 · Abstract (English)

LLMs can perform seemingly planning-intensive tasks, like writing coherent stories or functioning code, without explicitly verbalizing a plan; however, the extent to which they implicitly plan is unknown. In this paper, we define latent planning as occurring when LLMs possess internal planning representations that (1) cause the generation of a specific future token or concept, and (2) shape preceding context to license said future token or concept. We study the Qwen-3 family (0.6B-14B) on simple planning tasks, finding that latent planning ability increases with scale. Models that plan possess features that represent a planned-for word like "accountant", and cause them to output "an" rather than "a"; moreover, even the less-successful Qwen-3 4B-8B have nascent planning mechanisms. On the more complex task of completing rhyming couplets, we find that models often identify a rhyme ahead of time, but even large models seldom plan far ahead. However, we can elicit some planning that increases with scale when steering models towards planned words in prose. In sum, we offer a framework for measuring planning and mechanistic evidence of how models' planning abilities grow with scale.

大模型隐式规划规模效应生成机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。