用强化学习提升AI供应链代理的可靠性,降低决策波动风险。
Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management

- 通过系统级奖励微调共享大模型,优化决策策略。
- 最优模型可使成本降低67%,但存在显著决策不稳定性。
- 提出'代理牛鞭效应'并证明需改变策略而非简单平均输出。
本文基于MIT啤酒游戏,研究自主生成式AI代理在多层级供应链中的表现。识别出四个影响性能的推理时调控因素:模型选择、策略与安全机制、集中数据共享及提示工程。模型能力是主导因素:未经微调的推理模型已超越人类水平,优化后的模型相比人类团队可降低成本高达67%。然而,高均值表现掩盖了重大可靠性风险。我们提出‘代理牛鞭效应’,即运行间决策不稳定性在多层级系统中的放大现象。其中‘决策牛鞭’指由随机代理决策引发的订单波动,而非客户需求变化所致。即使需求路径固定,该不稳定性仍会跨设施瞬时放大或随时间累积。重复采样等常规测试时缓解方法无效,表明可靠性需依赖底层决策策略的改进。为此,我们提出基于组相对策略优化(GRPO)的强化学习后训练框架,使用系统级供应链奖励对共享基础大模型进行训练。实验显示,后训练显著减少尾部事件,抑制代理牛鞭效应,提升自主供应链代理的可靠性。
原文摘要 · Abstract (English)
This paper studies autonomous generative AI agents in multi-echelon supply chains using the MIT Beer Game. We identify four inference-time levers that shape performance: model selection, policies and guardrails, centralized data sharing, and prompt engineering. Model capability is the dominant factor: an out-of-the-box reasoning model exceeds human-level performance, and optimized reasoning models reduce costs by up to 67% relative to human teams. However, strong average performance masks substantial reliability risks. We introduce agent bullwhip: the amplification of run-to-run decision instability in autonomous multi-echelon systems. A central component is decision bullwhip, the portion of order variability generated by stochastic agent decisions rather than by changes in customer demand. We show that decision instability can amplify both across facilities at a fixed point in time and within the same facility over time, even when the demand path is held fixed. Repeated sampling, a natural test-time remedy, fails to meaningfully reduce this instability, suggesting that reliability requires changing the underlying decision policy rather than merely averaging over model outputs. To address this limitation, we propose a Group Relative Policy Optimization (GRPO)-based reinforcement-learning post-training framework that trains a shared base LLM using system-level supply-chain rewards. Post-training substantially reduces tail events, curtails agent bullwhip, and improves the reliability of autonomous supply-chain agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。