arXiv:2604.03605eess.SYcs.AI2026-04

提出结构感知的强化学习算法,优化异构任务下的多机器人调度。

Multi-Robot Multi-Queue Control via Exhaustive Assignment Actor-Critic Learning

  • 基于显式服务结构设计策略网络,强制执行耗尽式服务。
  • 在不同负载和异构到达率下,平均队列长度降低23%以上。
  • 适合动态环境中的实时多机器人任务分配场景。

我们研究具有非对称随机到达和切换延迟的多机器人、多队列系统的在线任务分配问题。采用离散时间建模:每个时段每个位置最多容纳一台机器人,处理任务消耗一个时段,跨位置移动需一时段延迟,各位置到达为独立的伯努利过程且到达率异质。基于此前的结构性结论——最优策略为耗尽型,我们构建折扣成本马尔可夫决策过程,并提出一种耗尽分配的演员-评论家架构,通过构造保证耗尽服务,仅学习空闲机器人的下一队列分配。与仅在对称条件下最优的耗尽服务最长(ESL)规则不同,该策略能适应到达率差异。在不同服务器-位置比、负载及异构到达配置下,所提策略始终低于ESL基线的折扣持有成本与平均队列长度,且在存在最优基准的实例中接近最优。结果表明,结构感知的演员-评论家方法为实时多机器人调度提供了有效途径。

原文摘要 · Abstract (English)

We study online task allocation for multi-robot, multi-queue systems with asymmetric stochastic arrivals and switching delays. We formulate the problem in discrete time: each location can host at most one robot per slot, servicing a task consumes one slot, switching between locations incurs a one-slot travel delay, and arrivals at locations are independent Bernoulli processes with heterogeneous rates. Building on our previous structural result that optimal policies are of exhaustive type, we formulate a discounted-cost Markov decision process and develop an exhaustive-assignment actor-critic policy architecture that enforces exhaustive service by construction and learns only the next-queue allocation for idle robots. Unlike the exhaustive-serve-longest (ESL) queue rule, whose optimality is known only under symmetry, the proposed policy adapts to asymmetry in arrival rates. Across different server-location ratios, loads, and asymmetric arrival profiles, the proposed policy consistently achieves lower discounted holding cost and smaller mean queue length than the ESL baseline, while remaining near-optimal on instances where an optimal benchmark is available. These results show that structure-aware actor-critic methods provide an effective approach for real-time multi-robot scheduling.

多机器人强化学习任务分配队列控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。