arXiv:2607.05683cs.LGcs.MA2026-07

用强化学习动态优化仓库机器人充电,提升订单完成率6%。

Deep Reinforcement Learning for Dynamic Battery Management of Autonomous Order Pickers

论文配图:Deep Reinforcement Learning for Dynamic Battery Management of Autonomous Order Pickers
图 1 · 摘自论文原文
  • 基于PPO的强化学习模型,自动决策充电站选择和充电时长。
  • 相比最强基线,订单完成率提升6%,充电总时长显著减少。
  • 适合物流仓储中需协调多机器人的动态充电场景。

仓库中自主移动机器人(AMRs)的电池充电是影响订单处理时效与吞吐量的关键挑战。本文针对随机订单到达下的动态充电问题,提出基于近端策略优化(PPO)的深度强化学习框架,适用于配备固定充电站的多区仓库。模型动态学习两个关键决策:充电站选择与最优充电时长,并显式考虑站点预期排队时间。大量数值实验表明,该方法相比先进DRL与传统启发式算法,订单完成率最高提升6%,同时大幅降低充电总耗时。模型在不同仓库配置与随机到达率下均表现稳健。最后,通过解析学习到的策略,揭示其优于基准方法的运营逻辑。

原文摘要 · Abstract (English)

Battery charging of Autonomous Mobile Robots (AMRs) in warehouses is a critical operational challenge that heavily impacts both order processing times and throughput. In this study, we address the dynamic AMR charging problem under stochastic order arrivals, where robots must learn optimal charging decisions. Traditional fixed-rule heuristics often prove suboptimal in dynamic environments and fail to account for multi-AMR coordination, leading to severe resource inefficiencies. To overcome these limitations, we propose a Proximal Policy Optimization (PPO)-based Deep Reinforcement Learning (DRL) framework designed for multi-block warehouses with fixed charging stations. Our model dynamically learns two key decisions: charging station selection and optimal charging duration, explicitly accounting for anticipated queuing times at the stations. Extensive numerical experiments benchmark the proposed model against state-of-the-art DRL and traditional heuristic approaches. Results demonstrate that our PPO framework increases order-completion rates by up to 6\% compared to the strongest baseline, while significantly reducing the total time dedicated to recharging operations. Furthermore, we validate the model's robustness across diverse warehouse configurations and stochastic arrival rates. Finally, we interpret the learned DRL policy, offering valuable operational insights into its superiority over standard benchmarks.

强化学习机器人调度仓储优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。