顶尖大模型首次在规划任务上超越传统算法,打破语言模型不会规划的固有认知。
Frontier Large Language Models Rival State-of-the-Art Planners
- 用最新大模型测试国际规划竞赛题,严格验证避免数据泄露。
- Gemini 3.1 Pro 解出245个任务(共360个),超过最强传统规划器。
- 模型代际对比显示能力持续提升,未来或可拓展至复杂推理场景。
多项重要研究曾表明,大语言模型无法可靠解决哪怕简单的规划任务。本文证明,最新一代前沿模型推翻了这一结论。我们在基于最新国际规划竞赛的挑战性规划任务上评估了三类前沿大模型,遵循严格评估标准:使用验证工具确认解的有效性,任务为新生成以避免数据污染,并与最先进经典规划器进行对比。在标准任务描述下,Gemini 3.1 Pro 解出245个任务(共360个),优于最强规划器基准(234个);GPT-5表现与基准相当。当所有语义信息被隐藏以测试纯符号规划能力时,性能下降但Gemini 3.1 Pro仍具竞争力。纵向对比显示,从GPT-3.5(零解)到GPT-5,模型能力呈显著上升趋势。前沿大模型或许终于具备规划能力,关键问题是这一能力能延伸多远。
原文摘要 · Abstract (English)
A series of influential studies established that large language models cannot reliably solve even simple planning tasks. We show that the latest generation of frontier models overturns this conclusion. We evaluate three families of frontier LLMs on a challenging set of planning tasks based on the most recent International Planning Competition following rigorous evaluation guidelines: solutions are verified with a validation tool, tasks are freshly created to avoid data contamination, and performance is compared against state-of-the-art classical planners. On standard task descriptions, Gemini 3.1 Pro outperforms the strongest planner baseline (245 vs. 234 solved tasks out of 360), while GPT-5 achieves comparable performance to the baselines. When all semantic information is obfuscated from the descriptions to test for pure symbolic planning, performance degrades but Gemini 3.1 Pro remains competitive with the strongest baselines. A longitudinal comparison across model generations -- from GPT-3.5, which solves zero tasks, to GPT-5 -- reveals a striking upward trajectory. Frontier LLMs might finally be able to plan; the question now is how far this capability will extend.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。