arXiv:2502.13443cs.RO2025-02ICRA被引 6

用在线掩码学习让机器人更智能地堆箱子,兼顾物理特性。

Physics-Aware Robotic Palletization with Online Masking Inference

  • 通过在线学习动态生成动作掩码,引导强化学习避开无效操作。
  • 在真实机器人上实现稳定堆叠,比现有方法提升堆叠成功率。
  • 适合需要考虑密度、刚性等物理特性的自动化仓储场景。

在现代仓库与物流管理中,高效规划箱体堆叠,尤其是在物品到达序列不可预测的在线场景下,仍是关键挑战。现有方案通常处理箱体尺寸差异,却忽略了其固有的物理属性(如密度、刚性),而这些属性对实际应用至关重要。本文采用强化学习(RL)解决此问题,通过动作空间掩码引导策略选择合法动作。不同于依赖难以评估的启发式稳定性判断的以往方法,本框架利用在线学习动态训练动作掩码,无需人工设计启发规则。大量实验表明,所提方法优于现有最先进水平。此外,我们在真实机器人堆垛机上部署了学习到的任务规划器,验证了其在实际运行环境中的可行性。

原文摘要 · Abstract (English)

The efficient planning of stacking boxes, especially in the online setting where the sequence of item arrivals is unpredictable, remains a critical challenge in modern warehouse and logistics management. Existing solutions often address box size variations, but overlook their intrinsic and physical properties, such as density and rigidity, which are crucial for real-world applications. We use reinforcement learning (RL) to solve this problem by employing action space masking to direct the RL policy toward valid actions. Unlike previous methods that rely on heuristic stability assessments which are difficult to assess in physical scenarios, our framework utilizes online learning to dynamically train the action space mask, eliminating the need for manual heuristic design. Extensive experiments demonstrate that our proposed method outperforms existing state-of-the-arts. Furthermore, we deploy our learned task planner in a real-world robotic palletizer, validating its practical applicability in operational settings.

机器人强化学习堆叠规划物理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。