arXiv:2504.09831stat.MLcs.AI2025-04被引 3

解决库存与定价中因需求被截断且依赖历史而难优化的难题。

Offline Dynamic Inventory and Pricing Strategy: Addressing Censored and Dependent Demand

  • 构建高阶马尔可夫决策过程,用连续截断次数建模需求依赖性
  • 提出两个新算法,基于离线数据学习最优定价与补货策略
  • 首次实现截断且相关需求下的数据驱动最优策略学习

本文研究离线序列特征定价与库存控制问题,其中当前需求依赖于历史需求水平,且超过库存量的需求将丢失。目标是利用包含过去价格、订货量、库存水平、协变量和截断销售量的离线数据集,估计能最大化长期利润的最优定价与库存控制策略。尽管无截断时可用马尔可夫决策过程(MDP)建模,但观测过程中存在需求截断,导致利润信息缺失、马尔可夫性失效及最优策略非平稳。为克服这些挑战,我们首先通过由连续截断次数决定的高阶MDP近似最优策略,最终转化为求解针对该问题定制的贝尔曼方程。受离线强化学习与生存分析启发,提出两种新颖的数据驱动算法求解贝尔曼方程,从而估计最优策略。进一步建立了有限样本后悔界以验证算法有效性。最后,数值实验展示了算法在估计最优策略方面的高效性。据我们所知,这是首个在具有截断与依赖需求的序列决策环境中实现数据驱动最优策略学习的方法。算法实现见 https://github.com/gundemkorel/Inventory_Pricing_Control

原文摘要 · Abstract (English)

In this paper, we study the offline sequential feature-based pricing and inventory control problem where the current demand depends on the past demand levels and any demand exceeding the available inventory is lost. Our goal is to leverage the offline dataset, consisting of past prices, ordering quantities, inventory levels, covariates, and censored sales levels, to estimate the optimal pricing and inventory control policy that maximizes long-term profit. While the underlying dynamic without censoring can be modeled by Markov decision process (MDP), the primary obstacle arises from the observed process where demand censoring is present, resulting in missing profit information, the failure of the Markov property, and a non-stationary optimal policy. To overcome these challenges, we first approximate the optimal policy by solving a high-order MDP characterized by the number of consecutive censoring instances, which ultimately boils down to solving a specialized Bellman equation tailored for this problem. Inspired by offline reinforcement learning and survival analysis, we propose two novel data-driven algorithms for solving these Bellman equations and, thus, estimate the optimal policy. Furthermore, we establish finite-sample regret bounds to validate the effectiveness of these algorithms. Finally, we conduct numerical experiments to demonstrate the efficacy of our algorithms in estimating the optimal policy. To the best of our knowledge, this is the first data-driven approach to learning optimal pricing and inventory control policies in a sequential decision-making environment characterized by censored and dependent demand. The implementations of the proposed algorithms are available at https://github.com/gundemkorel/Inventory_Pricing_Control

库存优化定价策略离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。