arXiv:2504.14773cs.AIcs.CL2025-04被引 9

构建首个大模型规划能力评估基准集,覆盖多场景测试。

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities

  • 系统梳理5类规划任务的现有基准,建立分类框架。
  • 提出适配不同算法的推荐基准组合,指导研究选型。
  • 为未来规划基准设计提供可参考的缺口分析。

规划是智能体与代理式AI的核心能力。例如,在预算内制定旅行计划的能力在科学和商业领域具有巨大潜力,且最优规划通常比临时方法更节省资源。然而,目前对现有规划基准的全面理解仍显不足,导致跨领域比较规划算法性能或为新场景选择合适算法面临挑战。本文系统考察了一系列规划基准,识别出算法开发中常用的任务测试平台,并指出潜在空白。这些基准被分为具身环境、网页导航、调度、游戏与谜题、日常任务自动化五类。研究推荐了各类算法适用的基准组合,并为未来基准建设提供了洞见。

原文摘要 · Abstract (English)

Planning is central to agents and agentic AI. The ability to plan, e.g., creating travel itineraries within a budget, holds immense potential in both scientific and commercial contexts. Moreover, optimal plans tend to require fewer resources compared to ad-hoc methods. To date, a comprehensive understanding of existing planning benchmarks appears to be lacking. Without it, comparing planning algorithms' performance across domains or selecting suitable algorithms for new scenarios remains challenging. In this paper, we examine a range of planning benchmarks to identify commonly used testbeds for algorithm development and highlight potential gaps. These benchmarks are categorized into embodied environments, web navigation, scheduling, games and puzzles, and everyday task automation. Our study recommends the most appropriate benchmarks for various algorithms and offers insights to guide future benchmark development.

大模型评估规划能力基准测试智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。