测试大模型对规划语言PDDL的理解与生成能力,发现表现参差不齐。
An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
- 在20个模型上评估零样本下解析、生成和推理PDDL的能力
- 部分模型在简单任务中表现良好,复杂规划仍存短板
- 适合研究AI规划或大模型推理能力的学者参考
近期大型语言模型(LLMs)在代码生成和链式思维推理方面展现出能力,为自动形式化规划任务奠定了基础。本研究评估了LLMs理解与生成人工智能规划中的核心表示语言——规划域定义语言(PDDL)的潜力。我们在涵盖7个主要家族的20个不同模型上开展全面分析,包括商用与开源模型。评估覆盖零样本下的PDDL解析、生成与推理能力。结果表明,尽管部分模型在处理PDDL时表现显著有效,但其他模型在需要精细规划知识的复杂场景中存在局限。这些发现揭示了LLMs在形式化规划任务中的前景与当前瓶颈,为智能驱动规划范式的应用提供了洞见,并指引未来研究方向。
原文摘要 · Abstract (English)
In recent advancements, large language models (LLMs) have exhibited proficiency in code generation and chain-of-thought reasoning, laying the groundwork for tackling automatic formal planning tasks. This study evaluates the potential of LLMs to understand and generate Planning Domain Definition Language (PDDL), an essential representation in artificial intelligence planning. We conduct an extensive analysis across 20 distinct models spanning 7 major LLM families, both commercial and open-source. Our comprehensive evaluation sheds light on the zero-shot LLM capabilities of parsing, generating, and reasoning with PDDL. Our findings indicate that while some models demonstrate notable effectiveness in handling PDDL, others pose limitations in more complex scenarios requiring nuanced planning knowledge. These results highlight the promise and current limitations of LLMs in formal planning tasks, offering insights into their application and guiding future efforts in AI-driven planning paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。