测试大工具生态中LLM智能体的长程规划能力,模拟真实环境下的工具失效与干扰。
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems

- 构建327个零售任务、1665个工具的交互式评测基准,评估智能体逐步发现并调用工具的能力。
- 在无干扰下GPT-5.4准确率达51.90%,严重干扰时骤降至11.36%,暴露规划脆弱性。
- 适合研究长程规划、工具调用与鲁棒性智能体的开发者与研究人员参考。
大型语言模型(LLM)智能体日益运行于复杂的工具生态系统中,真实任务需在长周期内发现相关工具、推断隐含子目标,并适应动态环境。然而,现有评测基准极少考察在工具可见性受限条件下的规划能力。为此,我们提出PlanBench-XL,一个包含327个零售任务和1,665个工具的交互式基准,用于检验智能体能否通过迭代检索可用工具,并调用它们以获取中间证据,从而推进至最终目标。该基准还引入可选阻塞机制,通过工具缺失、失败或干扰来模拟现实世界的不可预测性,迫使智能体在运行时检测路径中断并自适应调整。对十款主流LLM的实验表明,大规模工具规划仍具挑战:在无阻塞条件下,GPT-5.4准确率为51.90%,但在最严重阻塞条件下骤降至11.36%。进一步分析显示,当故障无明确错误信号,或恢复需更长替代工具路径时,智能体尤为脆弱。结果确立了PlanBench-XL作为诊断智能体规划失败的试验平台,并强调在复杂不完美工具环境中实现鲁棒自适应规划的必要性。
原文摘要 · Abstract (English)
LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons. However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility. To address this gap, we introduce PlanBench-XL, an interactive benchmark of 327 retail tasks over 1,665 tools that tests whether agents can iteratively retrieve usable tools, invoke them to uncover intermediate evidence for subsequent calls toward the final goal. PlanBench-XL further features an optional blocking mechanism that simulates real-world unpredictability through missing, failing, or distracting tool functions, forcing agents to detect disrupted paths and adapt at runtime. Experiments on ten leading LLMs show that massive-tool planning remains challenging: while GPT-5.4 achieves 51.90% accuracy in block-free settings, it collapses to 11.36% under the most severe blocking condition. Further analysis shows that agents are especially vulnerable when failures lack explicit error signals or when recovery requires longer alternative tool-use paths. These results establish PlanBench-XL as a testbed for diagnosing agentic planning failures and highlight the need for robust adaptive planning in long-horizon tasks with large, imperfect tool environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。