arXiv:2601.20856cs.AI2026-01被引 2

用推箱子游戏测试大模型的长程规划能力,发现超过25步就明显变差。

SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models

  • 设计推箱子简化版基准测试,专注评估长期规划能力
  • 超过25步时规划准确率显著下降,暴露模型瓶颈
  • 引入PDDL工具仅小幅提升效果,说明架构限制难靠调参突破

尽管大语言模型在复杂推理任务上的能力不断被检验,其长程规划能力仍缺乏系统研究。本文对前沿大推理模型(LRMs)的规划与长程推理能力进行了系统评估,提出基于推箱子谜题的新基准,刻意简化以隔离长期规划与状态保持的影响。结果显示,当解题所需步骤超过25步时,规划性能出现持续下降,表明存在本质性规划容量限制。实验表明,为模型配备规划领域定义语言(PDDL)解析、验证与求解工具可带来适度改进,暗示仅靠测试时扩展无法克服内在架构局限。

原文摘要 · Abstract (English)

Although the capabilities of large language models have been increasingly tested on complex reasoning tasks, their long-horizon planning abilities have not yet been extensively investigated. In this work, we provide a systematic assessment of the planning and long-horizon reasoning capabilities of state-of-the-art Large Reasoning Models (LRMs). We propose a novel benchmark based on Sokoban puzzles, intentionally simplified to isolate long-horizon planning from state persistence. Our findings reveal a consistent degradation in planning performance when more than 25 moves are required to reach the solution, suggesting a fundamental constraint on forward planning capacity. We show that equipping LRMs with Planning Domain Definition Language (PDDL) parsing, validation, and solving tools allows for modest improvements, suggesting inherent architectural limitations which might not be overcome by test-time scaling approaches alone.

长程规划大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。