arXiv:2602.05799math.OCcs.LG2026-02

动态需求下,自适应库存策略可有效应对未知变化并逼近最优解。

Non-Stationary Inventory Control with Lead Times

  • 设计自适应在线算法,在需求变化时优化基础库存策略。
  • 零提前期场景下,性能接近静态学习问题的理论最优水平。
  • 适合需求波动大、无历史数据的实时库存管理场景。

研究需求分布未知且随时间变化的单品类周期性库存控制问题。分析需求非平稳性对不同库存模型(含缺货回补或丢失销售)的学习性能影响,涵盖有无提前期的情形。针对每种设置,提出一种在基础库存策略类中优化的自适应在线算法,并建立相对于各时刻最优基础库存策略的动态后悔率性能保证。算法利用库存成本的凸性和单侧反馈结构,在需求被截断的情况下仍可进行反事实策略评估。在缺货回补系统和无提前期的丢失销售模型中,算法能适应未知需求变化,性能上至多相差对数因子,与已知静态学习问题的最优速率一致。而在存在正提前期的丢失销售系统中,需求截断与补货延迟共同限制了反事实评估,导致较弱的后悔界。通过仿真验证,所提方法显著优于现有非预言基准。

原文摘要 · Abstract (English)

We study non-stationary single-item, periodic-review inventory control problems in which the demand distribution is unknown and may change over time. We analyze how demand non-stationarity affects learning performance across inventory models, including systems with demand backlogging or lost-sales, both with and without lead times. For each setting, we propose an adaptive online algorithm that optimizes over the class of base-stock policies and establish performance guarantees in terms of dynamic regret relative to the optimal base-stock policy at each time step. The algorithms leverage the convexity and one-sided feedback structure of inventory costs to enable counterfactual policy evaluation despite demand censoring. In backlogging systems and lost-sales models with zero lead time, our algorithms adapt to unknown demand changes while matching, up to logarithmic factors, the rates known for the corresponding stationary learning problems. In lost-sales systems with positive lead times, the combination of demand censoring and delayed replenishment restricts counterfactual policy evaluation and leads to weaker regret guarantees. We complement the theoretical analysis with simulation results showing that our methods significantly outperform existing non-oracle benchmarks.

库存优化在线学习动态需求后悔分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。