arXiv:2409.14972cs.ROcs.AI2024-09被引 38

用深度强化学习让仓库机器人更智能避障,尤其能顾及行人舒适性。

Deep Reinforcement Learning-based Obstacle Avoidance for Robot Movement in Warehouse Environments

  • 通过行人角度网格与注意力机制,提升价值网络对行人的动态感知能力
  • 设计基于行人空间行为的奖励函数,避免机器人转向过快导致不适
  • 在复杂仓库仿真环境中验证了算法的有效性,适合智能仓储场景

当前大多数仓库环境货物堆积复杂,管理人与移动机器人路径交互频繁,传统移动机器人难以有效反馈针对货物和行人的正确避障策略。为使机器人在仓库环境中高效且友好地完成避障任务,本文提出一种基于深度强化学习的仓库环境移动机器人避障算法。首先,针对深度强化学习中价值函数网络学习能力不足的问题,基于行人交互改进价值函数网络:通过行人角度网格提取行人间交互信息,并利用注意力机制提取单个行人的时序特征,从而学习当前状态与历史轨迹状态的相对重要性及其联合影响,为后续多层感知机的学习提供支持。其次,基于行人空间行为设计强化学习奖励函数,对机器人转向角度变化过大的状态进行惩罚,以实现舒适的避障要求。最后,通过仿真实验验证了该算法在复杂仓库环境中的可行性和有效性。

原文摘要 · Abstract (English)

At present, in most warehouse environments, the accumulation of goods is complex, and the management personnel in the control of goods at the same time with the warehouse mobile robot trajectory interaction, the traditional mobile robot can not be very good on the goods and pedestrians to feed back the correct obstacle avoidance strategy, in order to control the mobile robot in the warehouse environment efficiently and friendly to complete the obstacle avoidance task, this paper proposes a deep reinforcement learning based on the warehouse environment, the mobile robot obstacle avoidance Algorithm. Firstly, for the insufficient learning ability of the value function network in the deep reinforcement learning algorithm, the value function network is improved based on the pedestrian interaction, the interaction information between pedestrians is extracted through the pedestrian angle grid, and the temporal features of individual pedestrians are extracted through the attention mechanism, so that we can learn to obtain the relative importance of the current state and the historical trajectory state as well as the joint impact on the robot's obstacle avoidance strategy, which provides an opportunity for the learning of multi-layer perceptual machines afterwards. Secondly, the reward function of reinforcement learning is designed based on the spatial behaviour of pedestrians, and the robot is punished for the state where the angle changes too much, so as to achieve the requirement of comfortable obstacle avoidance; Finally, the feasibility and effectiveness of the deep reinforcement learning-based mobile robot obstacle avoidance algorithm in the warehouse environment in the complex environment of the warehouse are verified through simulation experiments.

机器人避障强化学习仓库自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。