用强化学习优化仓储机器人充电策略,发现奖励设计影响学习效率与泛化能力。
Reinforcement Learning for AMR Charging Decisions: The Impact of Reward and Action Space Design
- 通过灵活奖励与动作空间设计,让机器人自主学习高效充电策略。
- 相比启发式方法,新方案平均服务时间减少18.3%,但开放设计收敛慢且不稳定。
- 适合关注机器人自主调度与强化学习工程实践的研究者。
我们提出一种新型强化学习(RL)框架,用于优化大规模堆叠仓库中自主移动机器人(AMR)的充电策略。该研究聚焦于不同奖励函数与动作空间配置对智能体性能的影响,涵盖从灵活自由到领域引导的设计。以启发式充电策略为基线,实验表明基于强化学习的方法在降低服务时间方面表现更优。研究发现:更开放的设计虽能自主发现高性能策略,但收敛时间长且学习过程不稳定;而有指导的设计则提升稳定性,但泛化能力受限。本文贡献包括:扩展开源仿真框架SLAPStack以支持充电策略建模;提出解决充电问题的新型强化学习设计;引入多种自适应基线启发式算法,并使用近端策略优化(PPO)算法系统评估不同设计配置,重点分析奖励设计的影响。
原文摘要 · Abstract (English)
We propose a novel reinforcement learning (RL) design to optimize the charging strategy for autonomous mobile robots in large-scale block stacking warehouses. RL design involves a wide array of choices that can mostly only be evaluated through lengthy experimentation. Our study focuses on how different reward and action space configurations, ranging from flexible setups to more guided, domain-informed design configurations, affect the agent performance. Using heuristic charging strategies as a baseline, we demonstrate the superiority of flexible, RL-based approaches in terms of service times. Furthermore, our findings highlight a trade-off: While more open-ended designs are able to discover well-performing strategies on their own, they may require longer convergence times and are less stable, whereas guided configurations lead to a more stable learning process but display a more limited generalization potential. Our contributions are threefold. First, we extend SLAPStack, an open-source, RL-compatible simulation-framework to accommodate charging strategies. Second, we introduce a novel RL design for tackling the charging strategy problem. Finally, we introduce several novel adaptive baseline heuristics and reproducibly evaluate the design using a Proximal Policy Optimization agent and varying different design configurations, with a focus on reward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。