arXiv:2507.02253cs.AI2025-07被引 3

自动生成可扩展的流程规划问题,评估大模型推理能力。

Scaling LLM Planning: NL2FLOW for Parametric Problem Generation and Rigorous Evaluation

  • 用结构化中间表示参数化生成流程规划问题。
  • 最优解生成率达69%,有效计划达86%。
  • 自然语言转结构化数据能显著提升成功率。

稳健的工作流组合对智能体性能至关重要,但大语言模型(LLM)在规划与推理方面的进展受限于可扩展评估数据的缺乏。本文提出NL2Flow,一个全自动的流程规划问题生成与评估流水线。该方法以结构化中间表示参数化生成问题,并将其转化为自然语言和形式化PDDL。在由NL2Flow生成的2296个低难度问题数据集上,评估了多个开源指令微调的LLM。结果表明,表现最佳的模型在生成有效计划方面达到86%成功率,在可解问题中生成最优计划的准确率为69%。回归分析显示,问题特征对计划生成的影响取决于模型和提示设计。重要的是,在符号规划前将自然语言问题转化为结构化JSON表示显著提升了成功率,表明神经符号融合的优势。这些发现强调了在系统规模扩大至更复杂任务时,理解错误来源的重要性。

原文摘要 · Abstract (English)

Robust workflow composition is critical for effective agent performance, yet progress in Large Language Model (LLM) planning and reasoning is hindered by a scarcity of scalable evaluation data. This work introduces NL2Flow, a fully automated pipeline for generating and evaluating workflow planning problems. NL2Flow generates problems parametrically in a structured intermediate representation, translating them into both natural language and formal PDDL. I evaluate several open-source, instruct-tuned LLMs on a dataset of 2296 low-difficulty problems generated by NL2Flow. Results demonstrate that the best-performing model achieved 86% success in generating valid plans and 69% in generating optimal plans (for solvable problems). Regression analysis shows that the influence of problem characteristics on plan generation is contingent on both model and prompt design. Importantly, translating natural language problems into a structured JSON representation prior to symbolic planning significantly improved success rates, suggesting a benefit from neuro-symbolic integration. These findings underscore the importance of understanding error sources within LLM reasoning as systems scale to more complex tasks. As LLM reasoning scales to increasingly complex problems, understanding the shifting bottlenecks and sources of error within these systems will be crucial.

大模型推理流程规划神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。