arXiv:2601.18137cs.AIcs.CL2026-01ACL被引 32

构建长时序智能体规划新基准,考验真实场景下的全局约束求解能力。

DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints

  • 设计多日旅行与多商品购物任务,融合主动信息获取与细粒度约束推理
  • 前沿大模型在复杂规划中仍表现不佳,暴露推理模式可靠性不足问题
  • 适合研究长周期智能体规划、工具使用与约束优化的学者和开发者

尽管智能体评估已转向长时序任务,多数基准仍侧重局部步骤推理,而非需要真正规划能力的全局约束优化(如时间与预算限制)。现有大语言模型规划基准也未能体现真实场景中主动信息获取与细粒度局部约束的特性。为此,我们提出DeepPlanning,一个面向实际长时序智能体规划的挑战性基准。其包含多日旅行规划和多商品购物任务,要求主动获取信息、进行细粒度局部推理以及全局约束优化。在DeepPlanning上的评估表明,即使最先进的代理型大模型在这些问题上仍表现困难,凸显了可靠显式推理模式与并行工具使用对实现更优效果-效率权衡的重要性。错误分析进一步指明了提升代理型大模型长时序规划能力的潜在方向。我们已开源代码与数据,以支持后续研究。

原文摘要 · Abstract (English)

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning ability. Meanwhile, existing LLM planning benchmarks underrepresent the active information gathering and fine-grained local constraints typical of real-world settings. To address this, we introduce DeepPlanning, a challenging benchmark for practical long-horizon agent planning. It features multi-day travel planning and multi-product shopping tasks that require proactive information acquisition, local constrained reasoning, and global constrained optimization. Evaluations on DeepPlanning show that even frontier agentic LLMs struggle with these problems, highlighting the importance of reliable explicit reasoning patterns and parallel tool use for achieving better effectiveness-efficiency trade-offs. Error analysis further points to promising directions for improving agentic LLMs over long planning horizons. We open-source the code and data to support future research.

智能体规划长时序任务约束优化大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。