用强化学习协同库存与推荐,提升企业整体利润。
Closing the Loop: Coordinating Inventory and Recommendation via Deep Reinforcement Learning on Multiple Timescales
- 分层设计多智能体强化学习,按职能分解策略并设定不同学习速度。
- 实验显示利润比孤立决策提高显著,且策略符合管理直觉。
- 适合需要跨部门协同优化的电商、零售等复杂业务场景。
有效的跨职能协同对提升企业整体利润至关重要,尤其在组织日益复杂和规模扩大的背景下。本文提出一个统一的多智能体强化学习框架,用于联合优化库存补货与个性化产品推荐。首先构建理论模型,刻画两者的内在关联,并推导出最优协同的分析基准,发现产品间及时间上的同步调整模式。基于此,设计一种多时标多智能体强化学习架构,按部门职能分解策略组件,根据任务复杂度和响应需求分配不同学习速率。该模型无关设计具备可扩展性和部署灵活性,多时标更新提升了收敛稳定性与异构决策的适应性。进一步证明了算法的渐近收敛性。大量仿真实验表明,相比孤立决策框架,本方法显著提升利润,且训练后的智能体行为与理论模型所得管理洞察高度一致。本工作为复杂商业环境中实现可扩展、可解释的跨职能协同提供了强化学习解决方案。
原文摘要 · Abstract (English)
Effective cross-functional coordination is essential for enhancing firm-wide profitability, particularly in the face of growing organizational complexity and scale. Recent advances in artificial intelligence, especially in reinforcement learning (RL), offer promising avenues to address this fundamental challenge. This paper proposes a unified multi-agent RL framework tailored for joint optimization across distinct functional modules, exemplified via coordinating inventory replenishment and personalized product recommendation. We first develop an integrated theoretical model to capture the intricate interplay between these functions and derive analytical benchmarks that characterize optimal coordination. The analysis reveals synchronized adjustment patterns across products and over time, highlighting the importance of coordinated decision-making. Leveraging these insights, we design a novel multi-timescale multi-agent RL architecture that decomposes policy components according to departmental functions and assigns distinct learning speeds based on task complexity and responsiveness. Our model-free multi-agent design improves scalability and deployment flexibility, while multi-timescale updates enhance convergence stability and adaptability across heterogeneous decisions. We further establish the asymptotic convergence of the proposed algorithm. Extensive simulation experiments demonstrate that the proposed approach significantly improves profitability relative to siloed decision-making frameworks, while the behaviors of the trained RL agents align closely with the managerial insights from our theoretical model. Taken together, this work provides a scalable, interpretable RL-based solution to enable effective cross-functional coordination in complex business settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。