arXiv:2409.13373cs.AIcs.CL2024-09被引 125

评测大模型能否规划,o1表现飞跃但仍不完善。

LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

  • 用PlanBench测试大模型与新型推理模型的规划能力
  • o1在该基准上性能远超其他模型,但未达饱和
  • 揭示部署前需关注准确性、效率与可靠性

规划能力被视为智能体的核心素质,自人工智能诞生以来始终是研究重点。尽管大型语言模型(LLMs)兴起后对其是否具备规划能力存有广泛讨论,但自2022年我们发布PlanBench这一可扩展基准以来,进展仍十分缓慢。OpenAI宣称其最新o1(Strawberry)模型专为突破传统自回归语言模型局限而设计,属于新型大型推理模型(LRM)。本文以此为契机,全面评估当前LLMs与新兴LRMs在PlanBench上的表现。结果表明,o1在该基准上实现了质的飞跃,显著领先于其他模型,但仍远未达到性能上限。这一提升也凸显出在实际部署前,必须审慎考虑其准确性、效率与可保证性等关键问题。

原文摘要 · Abstract (English)

The ability to plan a course of action that achieves a desired state of affairs has long been considered a core competence of intelligent agents and has been an integral part of AI research since its inception. With the advent of large language models (LLMs), there has been considerable interest in the question of whether or not they possess such planning abilities. PlanBench, an extensible benchmark we developed in 2022, soon after the release of GPT3, has remained an important tool for evaluating the planning abilities of LLMs. Despite the slew of new private and open source LLMs since GPT3, progress on this benchmark has been surprisingly slow. OpenAI claims that their recent o1 (Strawberry) model has been specifically constructed and trained to escape the normal limitations of autoregressive LLMs--making it a new kind of model: a Large Reasoning Model (LRM). Using this development as a catalyst, this paper takes a comprehensive look at how well current LLMs and new LRMs do on PlanBench. As we shall see, while o1's performance is a quantum improvement on the benchmark, outpacing the competition, it is still far from saturating it. This improvement also brings to the fore questions about accuracy, efficiency, and guarantees which must be considered before deploying such systems.

大模型评估推理能力规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。