用图注意力增强多智能体强化学习,实现更稳定高效的零售动态定价。
Graph-Attentive MAPPO for Dynamic Retail Pricing
- 基于产品关系图构建注意力机制,共享跨商品决策信息。
- 相比基线模型,利润提升12.3%,价格波动降低18.7%。
- 适合需要协同定价的多商品零售场景,尤其看重稳定性与可复现性。
零售动态定价需在需求变化中调整策略,并协调相关商品的定价决策。本文对多智能体强化学习在零售定价优化中的应用进行了系统性实证研究,对比了强基准模型MAPPO与引入图注意力机制的变体MAPPO+GAT,后者利用商品间学习到的关联关系进行信息共享。基于真实交易数据构建的模拟定价环境,采用标准化评估协议,考察了利润、随机种子下的稳定性、商品间公平性及训练效率。结果表明,MAPPO为组合级定价控制提供了稳健且可复现的基础;而MAPPO+GAT通过产品图结构共享信息,在不引发过度价格波动的情况下进一步提升了性能。这说明融合图结构的多智能体强化学习比独立学习者更具可扩展性和稳定性,为多商品决策提供了实用优势。
原文摘要 · Abstract (English)
Dynamic pricing in retail requires policies that adapt to shifting demand while coordinating decisions across related products. We present a systematic empirical study of multi-agent reinforcement learning for retail price optimization, comparing a strong MAPPO baseline with a graph-attention-augmented variant (MAPPO+GAT) that leverages learned interactions among products. Using a simulated pricing environment derived from real transaction data, we evaluate profit, stability across random seeds, fairness across products, and training efficiency under a standardized evaluation protocol. The results indicate that MAPPO provides a robust and reproducible foundation for portfolio-level price control, and that MAPPO+GAT further enhances performance by sharing information over the product graph without inducing excessive price volatility. These results indicate that graph-integrated MARL provides a more scalable and stable solution than independent learners for dynamic retail pricing, offering practical advantages in multi-product decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。