arXiv:2510.06475cs.AIcs.CL2025-10被引 5

用15类谜题评测大模型推理与规划能力,发现代码执行更高效但挑战更大。

PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles

  • 构建15类谜题基准,涵盖单双人、确定与随机场景。
  • 代码执行虽难但可扩展,推理模型在指令设置下表现更强。
  • 提供细粒度评估指标,助力模型优化与未来研究方向。

本工作探究基础模型在复杂动态环境中的推理与规划能力及其可扩展性。我们提出PuzzlePlex,一个通过多样化谜题评估这些能力的基准。PuzzlePlex包含15类谜题,涵盖不同难度的确定性与随机性游戏,以及单人和双人场景。该框架为每类游戏提供完整环境支持,并具备可扩展性,能随模型发展生成更复杂实例。我们还实现定制化博弈策略用于对比。基于此基准,我们开发细粒度评估指标,对前沿基础模型在指令驱动与代码执行两种设置下进行深度分析,并系统考察其扩展极限。结果表明,在指令设置中,推理模型表现更优;而代码执行虽更具挑战,但提供了可扩展且高效的替代方案。PuzzlePlex支持定向评估,推动基础模型在推理、规划与泛化方面持续改进。

原文摘要 · Abstract (English)

This work investigates the reasoning and planning capabilities of foundation models and their scalability in complex, dynamic environments. We introduce PuzzlePlex, a benchmark designed to assess these capabilities through a diverse set of puzzles. PuzzlePlex consists of 15 types of puzzles, including deterministic and stochastic games of varying difficulty, as well as single-player and two-player scenarios. The PuzzlePlex framework provides a comprehensive environment for each game, and supports extensibility to generate more challenging instances as foundation models evolve. Additionally, we implement customized game-playing strategies for comparison. Building on this benchmark, we develop fine-grained metrics to measure performance and conduct an in-depth analysis of frontier foundation models across two settings: instruction-based and code-based. Furthermore, we systematically investigate their scaling limits. Our findings show that reasoning models outperform others in instruction-based settings, while code-based execution presents greater challenges but offers a scalable and efficient alternative. PuzzlePlex enables targeted evaluation and guides future improvements in reasoning, planning, and generalization for foundation models.

推理能力基准测试规划模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。