揭示语言智能体规划能力不足的两大根本障碍
Revealing the Barriers of Language Agents in Planning
- 通过特征归因分析,发现约束和问题影响有限是关键瓶颈
- 即使最强模型o1在复杂任务上准确率仅15.6%
- 研究为提升规划智能体提供可解释性方向,适合对AI推理机制感兴趣者
自主规划自人工智能诞生以来就是持续追求的目标。早期规划智能体基于精心设计的问题求解器,可针对特定任务提供精确解,但缺乏泛化能力。大语言模型(LLMs)及其强大的推理能力重新激发了对自主规划的兴趣,使智能体能自动生成合理解决方案。然而,先前研究及我们的实验表明,当前语言智能体仍不具备人类级规划能力。即使最先进的推理模型OpenAI o1,在一个复杂现实规划基准测试中准确率也仅为15.6%。这提出了一个关键问题:是什么阻碍了语言智能体实现人类级规划?尽管已有研究指出智能体规划性能薄弱,但其深层原因、策略机制及其局限性仍理解不足。本文通过特征归因分析,识别出两个关键阻碍因素:约束作用有限与问题影响减弱。我们还发现,现有缓解策略虽有一定效果,但未能彻底解决问题,表明智能体距离人类智能仍有很长的路要走。
原文摘要 · Abstract (English)
Autonomous planning has been an ongoing pursuit since the inception of artificial intelligence. Based on curated problem solvers, early planning agents could deliver precise solutions for specific tasks but lacked generalization. The emergence of large language models (LLMs) and their powerful reasoning capabilities has reignited interest in autonomous planning by automatically generating reasonable solutions for given tasks. However, prior research and our experiments show that current language agents still lack human-level planning abilities. Even the state-of-the-art reasoning model, OpenAI o1, achieves only 15.6% on one of the complex real-world planning benchmarks. This highlights a critical question: What hinders language agents from achieving human-level planning? Although existing studies have highlighted weak performance in agent planning, the deeper underlying issues and the mechanisms and limitations of the strategies proposed to address them remain insufficiently understood. In this work, we apply the feature attribution study and identify two key factors that hinder agent planning: the limited role of constraints and the diminishing influence of questions. We also find that although current strategies help mitigate these challenges, they do not fully resolve them, indicating that agents still have a long way to go before reaching human-level intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。