arXiv:2601.07304cs.ROcs.AI2026-01

分拆导航与操作任务,让叉车在长时序多目标场景中更精准高效。

Heterogeneous Multi-Expert Reinforcement Learning for Long-Horizon Multi-Goal Tasks in Autonomous Forklifts

  • 用不同专家分别处理远距离导航和近距离操作,避免能力冲突。
  • 实验显示任务成功率94.2%,比基线提升31.7个百分点。
  • 适合需要高精度物料搬运的自动化仓储系统研发人员。

在非结构化仓库中实现自主移动操作,需平衡大规模导航效率与高精度物体交互。传统端到端学习方法难以兼顾这两个阶段的矛盾需求:导航依赖大空间鲁棒决策,而操作则需对局部细节高度敏感。单一网络同时学习两类目标常引发优化干扰,即提升一项任务会损害另一项。为此,我们提出专为自主叉车设计的异构多专家强化学习(HMER)框架。HMER将长时序任务分解为由语义任务规划器协调的专用子策略,分离宏观导航与微观操作,使各专家专注自身动作空间,避免干扰。规划器负责顺序调度专家,弥合任务规划与连续控制间的差距。此外,为解决探索稀疏问题,引入混合模仿-强化训练策略:利用专家示范初始化策略,再通过强化学习微调。Gazebo仿真结果表明,HMER显著优于串行与端到端基线,任务成功率达94.2%(基线62.5%),操作时间减少21.4%,放置误差保持在1.5厘米内,验证了其在精确物料搬运中的有效性。

原文摘要 · Abstract (English)

Autonomous mobile manipulation in unstructured warehouses requires a balance between efficient large-scale navigation and high-precision object interaction. Traditional end-to-end learning approaches often struggle to handle the conflicting demands of these distinct phases. Navigation relies on robust decision-making over large spaces, while manipulation needs high sensitivity to fine local details. Forcing a single network to learn these different objectives simultaneously often causes optimization interference, where improving one task degrades the other. To address these limitations, we propose a Heterogeneous Multi-Expert Reinforcement Learning (HMER) framework tailored for autonomous forklifts. HMER decomposes long-horizon tasks into specialized sub-policies controlled by a Semantic Task Planner. This structure separates macro-level navigation from micro-level manipulation, allowing each expert to focus on its specific action space without interference. The planner coordinates the sequential execution of these experts, bridging the gap between task planning and continuous control. Furthermore, to solve the problem of sparse exploration, we introduce a Hybrid Imitation-Reinforcement Training Strategy. This method uses expert demonstrations to initialize the policy and Reinforcement Learning for fine-tuning. Experiments in Gazebo simulations show that HMER significantly outperforms sequential and end-to-end baselines. Our method achieves a task success rate of 94.2\% (compared to 62.5\% for baselines), reduces operation time by 21.4\%, and maintains placement error within 1.5 cm, validating its efficacy for precise material handling.

强化学习自主叉车多专家任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。