构建可扩展的工业分拣强化学习环境,模拟真实产线并测试算法表现。
SortingEnv: An Extendable RL-Environment for an Industrial Sorting Process
- 基于数字孪生思想模拟分拣流程,支持速度调节与多模式分拣
- 对比PPO、DQN等算法在优化效率上优于传统规则策略
- 适合研究工业场景中智能体迁移性与环境演化的应用
我们提出一个新型强化学习(RL)环境,用于优化工业分拣系统并研究智能体在动态环境中的行为。该环境通过数字孪生理念模拟物料流动过程,包含皮带速度、占用率等操作参数。为反映现实挑战,引入新传感器和先进设备升级,提供两个版本:基础版仅支持离散皮带速度调整,高级版增加多种分拣模式及更丰富的物料组成观测。详细描述了两种环境的观测空间、状态更新机制与奖励函数设计。进一步评估了PPO、DQN、A2C等常用强化学习算法相较于经典规则基代理(RBA)的性能表现。该框架不仅有助于优化工业流程,还为研究智能体在演化环境中的行为与可迁移性提供了基础,揭示了模型在实际应用中的表现与意义。
原文摘要 · Abstract (English)
We present a novel reinforcement learning (RL) environment designed to both optimize industrial sorting systems and study agent behavior in evolving spaces. In simulating material flow within a sorting process our environment follows the idea of a digital twin, with operational parameters like belt speed and occupancy level. To reflect real-world challenges, we integrate common upgrades to industrial setups, like new sensors or advanced machinery. It thus includes two variants: a basic version focusing on discrete belt speed adjustments and an advanced version introducing multiple sorting modes and enhanced material composition observations. We detail the observation spaces, state update mechanisms, and reward functions for both environments. We further evaluate the efficiency of common RL algorithms like Proximal Policy Optimization (PPO), Deep-Q-Networks (DQN), and Advantage Actor Critic (A2C) in comparison to a classical rule-based agent (RBA). This framework not only aids in optimizing industrial processes but also provides a foundation for studying agent behavior and transferability in evolving environments, offering insights into model performance and practical implications for real-world RL applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。