arXiv:2603.16888cs.LGcs.AI2026-03被引 1

用多智能体强化学习优化动态定价,兼顾利润、稳定与公平。

Multi-Agent Reinforcement Learning for Dynamic Pricing: Balancing Profitability,Stability and Fairness

  • 采用MAPPO和MADDPG算法,在模拟零售市场中协同调价。
  • MAPPO平均收益最高且波动小,稳定性优于其他方法。
  • 适合关注定价公平性与系统稳定性的电商平台研究者。

竞争性零售市场中的动态定价需应对需求波动和对手行为变化。本文系统评估了多智能体强化学习(MARL)方法——MAPPO与MADDPG——在竞争环境下的动态价格优化表现。基于真实零售数据构建的模拟市场环境,我们将其与独立学习的IDDPG基线进行对比,评估利润、随机种子下的稳定性、公平性及训练效率。结果表明,MAPPO在平均收益上表现最优且方差最小,展现出高稳定性和可复现性;而MADDPG虽利润略低,但各智能体间利润分配最公平。研究证实,尤其在大规模场景下,MARL(尤其是MAPPO)是独立学习的可扩展、稳定的替代方案。

原文摘要 · Abstract (English)

Dynamic pricing in competitive retail markets requires strategies that adapt to fluctuating demand and competitor behavior. In this work, we present a systematic empirical evaluation of multi-agent reinforcement learning (MARL) approaches-specifically MAPPO and MADDPG-for dynamic price optimization under competition. Using a simulated marketplace environment derived from real-world retail data, we benchmark these algorithms against an Independent DDPG (IDDPG) baseline, a widely used independent learner in MARL literature. We evaluate profit performance, stability across random seeds, fairness, and training efficiency. Our results show that MAPPO consistently achieves the highest average returns with low variance, offering a stable and reproducible approach for competitive price optimization, while MADDPG achieves slightly lower profit but the fairest profit distribution among agents. These findings demonstrate that MARL methods-particularly MAPPO-provide a scalable and stable alternative to independent learning approaches for dynamic retail pricing.

动态定价多智能体强化学习零售

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。