arXiv:2506.11723cs.RO2025-06

用轻量强化学习让多机器人实时高效配送物料,速度提升百倍。

Dynamic Collaborative Material Distribution System for Intelligent Robots In Smart Manufacturing

  • 设计目标引导奖励函数的轻量DRL模型,实现快速决策。
  • 实验表明计算时间缩短至毫秒级,较枚举法提速100倍。
  • 模型可部署于物联网设备等低算力终端,适合智能产线应用。

多机器人协作已成为智能制造的关键环节,高效规划与管理对节能降本至关重要。本文针对智能制造中多机器人动态多源单目标(DMS-SD)路径规划问题,提出一种轻量级深度强化学习(DRL)方法。现有方法如枚举法虽能生成大量近优解,但无法利用历史经验;而基于轨迹信息的方法仅使用有限历史数据,导致大地图下计算耗时长,难以满足实时性需求。所提DRL方法通过设计目标引导奖励函数,可高效训练并快速收敛至最优解。实验表明,经训练的DRL模型将下一次移动的计算时间压缩至毫秒级,相较枚举法提速达100倍。该模型可部署于物联网设备、手机等低算力终端,仅需少量计算资源,适用于实际智能产线场景。

原文摘要 · Abstract (English)

The collaboration and interaction of multiple robots have become integral aspects of smart manufacturing. Effective planning and management play a crucial role in achieving energy savings and minimising overall costs. This paper addresses the real-time Dynamic Multiple Sources to Single Destination (DMS-SD) navigation problem, particularly with a material distribution case for multiple intelligent robots in smart manufacturing. Enumerated solutions, such as in \cite{xiao2022efficient}, tackle the problem by generating as many optimal or near-optimal solutions as possible but do not learn patterns from the previous experience, whereas the method in \cite{xiao2023collaborative} only uses limited information from the earlier trajectories. Consequently, these methods may take a considerable amount of time to compute results on large maps, rendering real-time operations impractical. To overcome this challenge, we propose a lightweight Deep Reinforcement Learning (DRL) method to address the DMS-SD problem. The proposed DRL method can be efficiently trained and rapidly converges to the optimal solution using the designed target-guided reward function. A well-trained DRL model significantly reduces the computation time for the next movement to a millisecond level, which improves the time up to 100 times in our experiments compared to the enumerated solutions. Moreover, the trained DRL model can be easily deployed on lightweight devices in smart manufacturing, such as Internet of Things devices and mobile phones, which only require limited computational resources.

多机器人强化学习智能制造实时调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。