用鲁棒强化学习优化仓库分拣口映射,抗动态包裹流入波动。
Distributionally Robust Multi-Agent Reinforcement Learning for Dynamic Chute Mapping
- 基于群体分布鲁棒优化,学习对不同流入场景都稳定的分拣策略。
- 模拟中平均降低80%包裹循环,显著提升系统稳定性。
- 适合处理物流调度中不确定性和季节性变化的场景。
在亚马逊机器人仓库中,目的地到分拣口的映射问题对高效包裹分拣至关重要。然而,频繁变动的包裹流入速率常导致包裹循环增加。为此,本文提出分布鲁棒多智能体强化学习(DRMARL)框架,学习对流入率对抗性变化具有鲁棒性的映射策略。具体而言,DRMARL采用群体分布鲁棒优化(DRO),使策略不仅在平均表现上优异,且在不同子群体(如不同季节或运行模式)的流入率下均表现良好。该方法结合一种基于上下文老虎机的最坏情况流入分布预测器,显著降低探索成本,提升学习效率与可扩展性。大量仿真表明,即使在不同流入分布条件下,DRMARL仍能实现稳健的分拣口映射,平均减少80%的包裹循环。
原文摘要 · Abstract (English)
In Amazon robotic warehouses, the destination-to-chute mapping problem is crucial for efficient package sorting. Often, however, this problem is complicated by uncertain and dynamic package induction rates, which can lead to increased package recirculation. To tackle this challenge, we introduce a Distributionally Robust Multi-Agent Reinforcement Learning (DRMARL) framework that learns a destination-to-chute mapping policy that is resilient to adversarial variations in induction rates. Specifically, DRMARL relies on group distributionally robust optimization (DRO) to learn a policy that performs well not only on average but also on each individual subpopulation of induction rates within the group that capture, for example, different seasonality or operation modes of the system. This approach is then combined with a novel contextual bandit-based predictor of the worst-case induction distribution for each state-action pair, significantly reducing the cost of exploration and thereby increasing the learning efficiency and scalability of our framework. Extensive simulations demonstrate that DRMARL achieves robust chute mapping in the presence of varying induction distributions, reducing package recirculation by an average of 80\% in the simulation scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。