用双代理强化学习解决价格与补货频率不一致的动态优化问题
Dual-Agent Deep Reinforcement Learning for Dynamic Pricing and Replenishment
- 设计快慢双代理框架,分别处理高频定价和低频补货
- 证明单期利润在价格与库存上均呈凹性,利于优化
- 结合决策树与DRL,在多产品场景中实现快速收敛
针对价格与补货决策频率不一致的动态定价与补货问题,本文研究了基于价格的泊松需求参数化模型下的复杂性。我们证明了单期利润函数在价格与库存域内均为凹函数,为优化提供理论基础。通过整合基于决策树的机器学习方法,利用全面市场数据增强需求预测。采用两尺度随机逼近算法,有效应对定价与补货之间的频率差异,确保收敛至局部最优。进一步引入深度强化学习(DRL),提出一种快-慢双代理DRL算法,其中两个代理分别负责价格与库存更新,且在不同时间尺度上进行。单产品与多产品场景的数值实验验证了该方法的有效性。
原文摘要 · Abstract (English)
We study the dynamic pricing and replenishment problems under inconsistent decision frequencies. Different from the traditional demand assumption, the discreteness of demand and the parameter within the Poisson distribution as a function of price introduce complexity into analyzing the problem property. We demonstrate the concavity of the single-period profit function with respect to product price and inventory within their respective domains. The demand model is enhanced by integrating a decision tree-based machine learning approach, trained on comprehensive market data. Employing a two-timescale stochastic approximation scheme, we address the discrepancies in decision frequencies between pricing and replenishment, ensuring convergence to local optimum. We further refine our methodology by incorporating deep reinforcement learning (DRL) techniques and propose a fast-slow dual-agent DRL algorithm. In this approach, two agents handle pricing and inventory and are updated on different scales. Numerical results from both single and multiple products scenarios validate the effectiveness of our methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。