arXiv:2605.08758cs.ROcs.AI2026-05

提出统一框架,让分拣机器人系统更智能高效

Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems

论文配图:Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems
图 1 · 摘自论文原文
  • 融合组合优化与多智能体强化学习,统筹订单、料箱和机器人决策
  • 小规模下近似最优,大规模下减少8-12%料箱移动量
  • 适合电商与工业物流中的自动化分拣系统部署

随着电商发展和小批量生产兴起,成品、半成品及原材料的物流单元尺寸持续缩小,料箱正逐步取代托盘成为主要搬运与存储容器。这一趋势推动料箱处理机器人系统成为自动化分拣中心的核心。料箱处理系统的订单履行决策具有订单-料箱-机器人顺序决策特性。现有研究多针对特定系统设计决策机制,难以泛化或迁移。本文提出面向料箱处理机器人系统的全尺度学习型顺序决策框架(OLSF-TRS),结合结构化组合优化与多智能体强化学习,协调订单、料箱与机器人的决策。在小规模系统中,OLSF-TRS 在两种不同配置下平均最优性差距低于3.5%,接近最优。在大规模场景中,相较于启发式基线,其料箱总移动量减少8%-12%,相比最先进规则方法降低超30%,同时保持实时响应能力。这些改进带来显著运营效益,包括成本下降、能耗降低与吞吐量稳定性提升。该框架为广泛应用的料箱处理机器人系统提供高效统一的订单履行决策方案,支持电商与工业物流高质量履约。

原文摘要 · Abstract (English)

Driven by the rapid expansion of e-commerce and small-batch production, the size of the intralogistics load unit of finished goods, semi-finished goods and raw materials is steadily shrinking. Totes are gradually replacing pallets as the primary handling and storage container. This shift has propelled tote-handling robotic systems to the forefront of automation order fulfillment centers. The order-fulfillment decisions of tote-handling robotic systems share a common order-tote-robot sequential decision-making nature. Existing studies primarily focus on decision mechanisms tailored to particular systems, making it difficult to generalize or transfer them to other contexts. We propose an Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems (OLSF-TRS), a generalized and scalable sequential decision framework that combines structured combinatorial optimization with multi-agent reinforcement learning to coordinate order,tote, and robot decisions. On small-scale tote-handling robotic systems, OLSF-TRS achieves near-optimal performance with average optimality gaps below 3.5% across two distinct system configurations. In large-scale scenarios, OLSF-TRS consistently outperforms heuristic baselines across two different system types, reducing total tote movements by 8-12% and over 30% compared to SOTA rule-based approaches, while maintaining real-time responsiveness. These improvements translate into tangible operational benefits, including cost reduction, lower energy consumption, and enhanced throughput stability. The proposed framework delivers an efficient and unified order fulfillment decision-making framework for widely deployed tote-handling robotic systems,supporting high-quality order fulfillment in both e-commerce and industrial logistics sectors.

机器人调度智能决策物流自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。