arXiv:2607.11197cs.AI2026-07

发现大模型规划能力有两类,影响表现的关键不是任务难易,而是能力类型。

What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities

论文配图:What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
图 1 · 摘自论文原文
  • 用多维项目反应模型分析大模型规划能力,识别出两种核心能力。
  • 操作推理随模型规模和推理长度提升,结构枚举则变化不大。
  • 适合关注模型能力细节、评估不同规划任务的科研人员阅读。

当大模型在规划任务中表现不均时,通常归因于任务难度差异。我们认为这一解释不完整,因为任务间差异可能源于潜藏的两种不同规划能力,而非单一能力谱系。我们在ACPBench-Hard上评估多个大模型家族,在不同推理预算下进行测试,并应用多维项目反应理论模型揭示大模型规划的潜在能力结构。分析显示两个主维度:操作推理(评估局部动作适用性和即时状态转移)与结构枚举(推理目标可达性与关键点结构)。操作推理随模型规模扩大和推理长度增加而提升,而结构枚举则相对稳定。研究呼吁从能力层面评估大模型规划,关注哪些能力提升、在何种条件下以及为何提升。

原文摘要 · Abstract (English)

When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete, as task-level variation may reflect distinct latent planning competencies rather than differences along a single ability spectrum. We study this question on ACPBench-Hard by evaluating multiple LLM families under varying test-time reasoning budgets and applying a multidimensional item response theory model to uncover the latent competency structure underlying LLM planning. The analysis reveals two principal dimensions that shape planning performance: operational reasoning, the ability to evaluate local action applicability and immediate state transitions, and structural enumeration, the ability to reason about goal reachability and landmark structure. Operational reasoning improving under model scaling and longer reasoning traces, while structural enumeration remains comparatively insensitive. Our findings motivate competency-level evaluation of LLM planning, shifting the focus from whether models improve overall to which planning competencies improve, under what conditions, and why.

大模型规划能力认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。