arXiv:2602.12631cs.AIcs.HC2026-02被引 3

将人类、大模型与运筹学方法结合,提升库存决策效果。

AI Agents for Inventory Control: Human-LLM-OR Complementarity

  • 用大模型+运筹学算法协同决策,互补而非替代。
  • 在1000多个真实与合成场景中,联合方案利润更高。
  • 人机协作让多数个体受益,且有理论下限支撑。

库存控制是运筹学中的核心问题,传统上依赖理论严谨的运筹学(OR)算法。然而,这些算法常依赖严格假设,在需求分布变化或上下文信息缺失时表现不佳。近年来大语言模型(LLM)的发展催生了能灵活推理并整合丰富上下文信号的AI代理,但如何将其融入传统决策流程仍不明确。本文研究在多周期库存控制场景中,运筹学算法、大模型与人类如何交互并相互补。我们构建了InventoryBench基准,包含超过1000个库存实例,涵盖合成与真实需求数据,用于在需求波动、季节性和不确定提前期等条件下测试决策规则。实验表明,经运筹学增强的大模型方法优于单独使用任一方法,证明三者具有互补性。进一步通过课堂控制实验,将大模型建议嵌入人机协同决策流程,发现平均而言,人机团队利润高于单独的人类或AI代理。除整体提升外,我们形式化了个体内互补效应,并推导出不依赖分布的个体受益比例下界;实证显示该比例显著。

原文摘要 · Abstract (English)

Inventory control is a fundamental operations problem in which ordering decisions are traditionally guided by theoretically grounded operations research (OR) algorithms. However, such algorithms often rely on rigid modeling assumptions and can perform poorly when demand distributions shift or relevant contextual information is unavailable. Recent advances in large language models (LLMs) have generated interest in AI agents that can reason flexibly and incorporate rich contextual signals, but it remains unclear how best to incorporate LLM-based methods into traditional decision-making pipelines. We study how OR algorithms, LLMs, and humans can interact and complement each other in a multi-period inventory control setting. We construct InventoryBench, a benchmark of over 1,000 inventory instances spanning both synthetic and real-world demand data, designed to stress-test decision rules under demand shifts, seasonality, and uncertain lead times. Through this benchmark, we find that OR-augmented LLM methods outperform either method in isolation, suggesting that these methods are complementary rather than substitutes. We further investigate the role of humans through a controlled classroom experiment that embeds LLM recommendations into a human-in-the-loop decision pipeline. Contrary to prior findings that human-AI collaboration can degrade performance, we show that, on average, human-AI teams achieve higher profits than either humans or AI agents operating alone. Beyond this population-level finding, we formalize an individual-level complementarity effect and derive a distribution-free lower bound on the fraction of individuals who benefit from AI collaboration; empirically, we find this fraction to be substantial.

库存优化人机协作大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。