评测大模型在真实供应链中的多步决策能力,发现现有模型可靠性不足。
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
- 构建统一基准,评估模型在供应链流程中的长时序工具调用能力
- 多模型测试显示执行可靠性差距显著,最高仅达62%准确率
- 提出无需SOP的自动生成执行流程框架,提升工具调用一致性
大型语言模型(LLMs)在复杂推理和基于工具的决策中展现出潜力,推动其在真实供应链管理中的应用。然而,供应链工作流需要基于领域特定规程的可靠长周期、多步骤编排,这对当前模型仍是挑战。为此,我们引入SupChain-Bench,一个统一的真实世界基准,评估模型在标准操作程序(SOPs)基础上的供应链领域知识与长时序工具编排能力。实验揭示了各模型在执行可靠性上的巨大差距。我们进一步提出SupChain-ReAct,一种无需SOP的框架,可自主合成可执行的工具使用流程,在工具调用性能上达到最强且最一致表现。本工作建立了研究真实运营场景中可靠长周期编排的基准,并指出了大模型供应链代理仍有显著改进空间。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in complex reasoning and tool-based decision making, motivating their application to real-world supply chain management. However, supply chain workflows require reliable long-horizon, multi-step orchestration grounded in domain-specific procedures, which remains challenging for current models. To systematically evaluate LLM performance in this setting, we introduce SupChain-Bench, a unified real-world benchmark that assesses both supply chain domain knowledge and long-horizon tool-based orchestration grounded in standard operating procedures (SOPs). Our experiments reveal substantial gaps in execution reliability across models. We further propose SupChain-ReAct, an SOP-free framework that autonomously synthesizes executable procedures for tool use, achieving the strongest and most consistent tool-calling performance. Our work establishes a principled benchmark for studying reliable long-horizon orchestration in real-world operational settings and highlights significant room for improvement in LLM-based supply chain agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。