arXiv:2608.14096cs.LGmath.OC2026-08

动态调整库存分配策略,应对需求不确定性

Resource-Adaptive Primal-Dual Learning for One-Warehouse Multi-Store Systems with Censored Demand

  • 根据剩余库存实时更新分配目标与决策变量
  • 实现对数级期望后悔值,优于现有方法的根号阶表现
  • 适合库存随时间消耗的多门店系统优化

单仓库多门店(OWMS)系统是库存网络中的基础模型,其中不可补货的仓库需在多门店间分配共享库存。现有学习策略基于初始平均资源率设定固定目标,无法随实际销售变化调整剩余资源的分配基准。本文提出资源自适应原-对偶学习框架,通过追踪随剩余资源演变的原-对偶求解路径,利用被截断的销售数据作为梯度估计,实时更新目标分配与对偶变量。分析结合期望销量几何结构与移动目标论证,证明其达到对数级期望后悔值,优于现有方法的平方根阶保证。数值实验表明,该框架在不同周期长度与库存条件下均具良好有限时间性能。

原文摘要 · Abstract (English)

The one-warehouse multi-store (OWMS) system is a fundamental inventory network in which a nonreplenishable warehouse allocates shared stock across multiple stores over time. Existing OWMS learning policies are built around a fixed target calibrated to the initial average resource rate, but such a fixed-target architecture cannot re-center after realized sales change the remaining resource available per future period. We develop Resource-Adaptive Primal-Dual Learning, a new learning framework that tracks the primal-dual resolving path with censored demand as the remaining-resource state evolves. In each period, the current resource rate indexes the target store allocations and dual variable, while censored sales provide gradient estimates for updating both. The analysis combines expected-sales geometry with a moving-target argument to yield logarithmic expected regret, improving on the state-of-the-art square-root-order guarantees of existing OWMS learning policies. The underlying design and analytical ideas may inform other online learning problems with depleting shared resources. Numerical experiments further demonstrate good finite-horizon performance of a practical variant across different horizon lengths and inventory regimes.

库存优化在线学习资源分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。