arXiv:2604.10953cs.RO2026-04

用扩散强化学习优化3D装箱,提升装入数量。

Diffusion Reinforcement Learning Based Online 3D Bin Packing Spatial Strategy Optimization

  • 构建马尔可夫决策链建模装箱过程
  • 相比顶尖方法平均多装入23%的物品
  • 适合复杂动态物流场景部署

在线3D装箱问题在物流、仓储和智能制造中至关重要,现有解决方案逐渐转向深度强化学习(DRL),但面临样本效率低的问题。本文提出一种基于扩散强化学习的算法,采用马尔可夫决策链建模装箱过程,使用高度图进行状态表示,并设计基于扩散模型的策略网络。实验表明,该方法显著提升了平均装入物品数量,优于当前最先进的DRL方法,在复杂在线场景中具有良好的应用潜力。

原文摘要 · Abstract (English)

The online 3D bin packing problem is important in logistics, warehousing and intelligent manufacturing, with solutions shifting to deep reinforcement learning (DRL) which faces challenges like low sample efficiency. This paper proposes a diffusion reinforcement learning-based algorithm, using a Markov decision chain for packing modeling, height map-based state representation and a diffusion model-based actor network. Experiments show it significantly improves the average number of packed items compared to state-of-the-art DRL methods, with excellent application potential in complex online scenarios.

3D装箱强化学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。