arXiv:2602.24182cs.LG2026-02

用多目标强化学习优化人机协同仓库的货箱分配,兼顾效率与资源约束。

Multi-Objective Reinforcement Learning for Large-Scale Tote Allocation in Human-Robot Collaborative Fulfillment Centers

  • 基于零和博弈的最优响应机制,实现多目标权衡的策略学习
  • 单个策略同时满足所有约束,仿真中显著提升空间与资源利用率
  • 提出新理论框架解决误差累积问题,适合工业级复杂决策场景

在基于容器的履约中心中,优化整合流程需权衡处理速度、资源使用与空间利用等多重目标,并遵守各类实际运营约束。该过程通过人机工作站协作移动货物,为入站库存腾出空间并提高容器利用率。本文将此问题建模为高维状态空间下的大规模多目标强化学习(MORL)任务,借鉴近期关于约束强化学习中最小最大策略学习的理论进展,采用最优响应与无悔动态方法求解。在真实仓库仿真中的策略评估表明,该方法能有效平衡多个目标,且实验观察到单一策略可同时满足所有约束,即使理论上不保证。此外,我们提出一个理论框架以应对误差累积问题——时间平均解出现振荡行为,本方法返回的单次迭代解其拉格朗日值接近博弈的最小最大值。结果证明MORL在解决大规模工业系统复杂决策问题上的潜力。

原文摘要 · Abstract (English)

Optimizing the consolidation process in container-based fulfillment centers requires trading off competing objectives such as processing speed, resource usage, and space utilization while adhering to a range of real-world operational constraints. This process involves moving items between containers via a combination of human and robotic workstations to free up space for inbound inventory and increase container utilization. We formulate this problem as a large-scale Multi-Objective Reinforcement Learning (MORL) task with high-dimensional state spaces and dynamic system behavior. Our method builds on recent theoretical advances in solving constrained RL problems via best-response and no-regret dynamics in zero-sum games, enabling principled minimax policy learning. Policy evaluation on realistic warehouse simulations shows that our approach effectively trades off objectives, and we empirically observe that it learns a single policy that simultaneously satisfies all constraints, even if this is not theoretically guaranteed. We further introduce a theoretical framework to handle the problem of error cancellation, where time-averaged solutions display oscillatory behavior. This method returns a single iterate whose Lagrangian value is close to the minimax value of the game. These results demonstrate the promise of MORL in solving complex, high-impact decision-making problems in large-scale industrial systems.

强化学习多目标优化仓储自动化人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。