arXiv:2508.12782cs.AI2025-08被引 6

评测大模型在虚拟世界中执行数百上千步规划的能力

HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds

  • 构建复杂游戏世界,要求模型端到端完成多步骤规划
  • 25个主流模型中无一能稳定解决最难题目,差距显著
  • 适合研究长时序推理与自主决策的开发者和学者

大型语言模型在数学和代码生成等逐步推理任务中表现良好,但在真实约束下进行鲁棒的长时程规划能力仍缺乏有效评估。现有规划基准多依赖抽象领域或交互反馈,掩盖了端到端规划失败和可行性错误。我们提出HeroBench,一个用于评估复杂类RPG虚拟世界中长时程、分层规划与结构化推理的基准。任务要求模型选择数值可行的装备,推理多层级制造与资源依赖关系,并以单一端到端计划执行数百至数千步操作。HeroBench融合符号规划、数值战斗模拟、空间推理与资源管理,支持可扩展难度与对抗性干扰。通过仿真评估可执行计划,提供成功指标与细粒度进展度量,以及详细的失败模式分析。对25个顶尖LLM的评估显示性能差异巨大,远超传统推理基准;尽管推理模型表现更优,但无一能可靠解决最难题目,凸显长时程自主规划的持续挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) perform well on step-by-step reasoning benchmarks such as mathematics and code generation, yet their ability to carry out robust long-horizon planning under realistic constraints remains insufficiently evaluated. Existing planning benchmarks often rely on abstract domains or interactive feedback, obscuring end-to-end planning failures and feasibility errors. We introduce HeroBench, a benchmark for evaluating long-horizon, hierarchical planning and structured reasoning in a complex RPG-inspired virtual world. Tasks require models to select numerically feasible equipment, reason over multi-level crafting and resource dependencies, and execute hundreds to thousands of actions as a single end-to-end plan. HeroBench integrates symbolic planning, numeric combat simulation, spatial reasoning, and resource management, while supporting scalable difficulty and adversarial distractors. HeroBench evaluates executable plans through simulation, enabling both success-based and fine-grained progress metrics, as well as detailed failure mode analysis. An evaluation of 25 state-of-the-art LLMs reveals large performance disparities rarely observed in conventional reasoning benchmarks. While reasoning models perform substantially better, no model reliably solves the hardest tasks, highlighting persistent challenges in long-horizon autonomous planning.

长时序规划虚拟世界推理评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。