arXiv:2503.17573cs.LG2025-03中稿 · presentation at IC…

用深度强化学习优化有高度限制的二维+一维包装问题

Optimizing 2D+1 Packing in Constrained Environments Using Deep Reinforcement Learning

  • 基于DRL设计可选位置与物料类型的多离散动作策略
  • PPO方法在资源利用率上优于经典启发式算法MaxRect-BL
  • 适用于航空航天复合材料等工业场景的复杂包装优化

本文提出一种基于深度强化学习(DRL)的新方法,解决带有空间约束的2D+1包装问题。该问题为传统2D包装的延伸,增加了高度维度的限制。为此,我们基于OpenAI Gym框架开发了一个仿真器,用于高效模拟矩形物料在两个板面上的放置过程,并支持多离散动作,可选择任意板面位置及待放置物料类型。采用两种DRL算法(PPO与A2C)学习包装策略,并与知名启发式基线方法MaxRect-BL进行对比。实验表明,基于PPO的方法在复杂包装任务中表现优异,展现出在航空航天复合材料制造等工业应用中优化资源利用的巨大潜力。

原文摘要 · Abstract (English)

This paper proposes a novel approach based on deep reinforcement learning (DRL) for the 2D+1 packing problem with spatial constraints. This problem is an extension of the traditional 2D packing problem, incorporating an additional constraint on the height dimension. Therefore, a simulator using the OpenAI Gym framework has been developed to efficiently simulate the packing of rectangular pieces onto two boards with height constraints. Furthermore, the simulator supports multidiscrete actions, enabling the selection of a position on either board and the type of piece to place. Finally, two DRL-based methods (Proximal Policy Optimization -- PPO and the Advantage Actor-Critic -- A2C) have been employed to learn a packing strategy and demonstrate its performance compared to a well-known heuristic baseline (MaxRect-BL). In the experiments carried out, the PPO-based approach proved to be a good solution for solving complex packaging problems and highlighted its potential to optimize resource utilization in various industrial applications, such as the manufacturing of aerospace composites.

强化学习包装优化工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。