用多智能体强化学习优化供应链动态定价,模拟真实市场博弈行为。
Multi-Agent Reinforcement Learning for Dynamic Pricing in Supply Chains: Benchmarking Strategic Agent Behaviours under Realistically Simulated Market Conditions
- 设计基于真实电商数据的多智能体仿真环境,对比三种MARL算法表现。
- MADDPG在公平性(0.8819)和价格稳定间取得平衡,竞争性更强。
- 相比静态规则,MARL能激发真实市场中的策略互动,适合研究定价博弈。
本研究探讨多智能体强化学习(MARL)在供应链动态定价中的应用,针对传统ERP系统依赖静态规则、忽略市场主体间策略互动的问题。通过结合真实电商业务数据与LightGBM需求预测模型,构建仿真环境,评估MADDPG、MADQN和QMIX三种MARL算法相对于静态规则基线的表现。结果表明,规则基线实现近乎完美的公平性(Jain指数:0.9896)和最高价格稳定性(波动率:0.024),但缺乏竞争动态;其中MADQN表现出最强攻击性,波动率高且公平性最低(0.5844);而MADDPG在保持较高公平性(0.8819)的同时,支持市场竞争(份额波动率:9.5个百分点),展现出更优的综合性能。研究证实MARL可催生传统规则无法捕捉的策略行为,为动态定价系统发展提供新思路。
原文摘要 · Abstract (English)
This study investigates how Multi-Agent Reinforcement Learning (MARL) can improve dynamic pricing strategies in supply chains, particularly in contexts where traditional ERP systems rely on static, rule-based approaches that overlook strategic interactions among market actors. While recent research has applied reinforcement learning to pricing, most implementations remain single-agent and fail to model the interdependent nature of real-world supply chains. This study addresses that gap by evaluating the performance of three MARL algorithms: MADDPG, MADQN, and QMIX against static rule-based baselines, within a simulated environment informed by real e-commerce transaction data and a LightGBM demand prediction model. Results show that rule-based agents achieve near-perfect fairness (Jain's Index: 0.9896) and the highest price stability (volatility: 0.024), but they fully lack competitive dynamics. Among MARL agents, MADQN exhibits the most aggressive pricing behaviour, with the highest volatility and the lowest fairness (0.5844). MADDPG provides a more balanced approach, supporting market competition (share volatility: 9.5 pp) while maintaining relatively high fairness (0.8819) and stable pricing. These findings suggest that MARL introduces emergent strategic behaviour not captured by static pricing rules and may inform future developments in dynamic pricing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。