用离线强化学习优化仓库SLAM吞吐量,提升效率同时保持系统稳定。
Offline Reinforcement Learning for Warehouse SLAM Throughput Control

- 基于历史数据的智能调控策略,动态平衡吞吐量与下游稳定性。
- CQL算法使系统健康度提升22.97%,平均限流时长减少3.18%。
- 适用于大规模仓储场景,可适配多种离线强化学习方法。
本文提出一种面向仓库履约环境的离线强化学习(RL)框架,用于优化SLAM(Scan/Label/Apply/Manifest)吞吐量控制。SLAM吞吐量直接影响系统拥堵与运营效率。所提方法通过智能调节限流行为,动态推荐吞吐量设置,在最大化吞吐量的同时保障下游稳定性。框架包含历史感知的状态表示、考虑延迟影响的动作空间抽象,以及融合上下游运营指标的奖励函数。该方法具备算法无关性,支持在统一架构下集成多种离线RL算法。我们使用来自大型仓库的真实脱敏运营日志,离线训练了三种前沿离线RL算法。通过多方法评估策略验证性能,包括基于回归模型的即时奖励估计、长周期Fitted Q Evaluation(FQE),以及基于Deep Koopman动力学的模型评估。实验结果表明,CQL策略表现最优,系统健康度提升22.97%,平均限流时长减少3.18%。这些发现证明了离线强化学习在安全、可扩展的仓库吞吐量优化中的潜力。
原文摘要 · Abstract (English)
We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment. SLAM (Scan/Label/Apply/Manifest) throughput directly influences system congestion and operational efficiency. Our RL-based control approach dynamically recommends SLAM throughput settings that adaptively balance throughput maximization with downstream stability through intelligent adjustment of throttling behavior. We include a history-informed state representation, action space abstraction for delayed-impact control, and a reward function that captures both upstream and downstream operational metrics. Our approach is algorithm-agnostic, enabling integration of multiple offline RL methods under a unified architecture. We instantiate our framework with three state-of-the-art offline RL algorithms, and trained the models offline using de-identified historical operational logs from a large-scale warehouse. Policy performance is evaluated using a comprehensive multi-method strategy. These include model-free approaches including immediate reward estimation via regression models and long-horizon Fitted Q Evaluation (FQE), as well as model-based Deep Koopman dynamics evaluation. Empirical results reveal that the CQL policy consistently outperforms alternatives, improving system health by 22.97% and reducing average throttling duration by 3.18%. These findings demonstrate the potential of offline RL for safe and scalable warehouse throughput control optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。