基于预测的动态调度框架,提升异构机器人系统在需求变化下的服务效率。
Model-Based Reinforcement Learning for Heterogeneous Multi-Robot Task Assignment Under Distribution Shifts

- 融合历史预测与实时反馈,动态生成任务分配策略
- 实测等待时间减少,尾部延迟改善显著,接近全服务率
- 适合医疗等高动态场景的机器人调度,尤其对预测偏差敏感
异构多机器人服务系统需将请求分配给兼容机器人,制定可行调度并在线响应新任务。历史数据可预判需求,但过度依赖不准确预测会降低性能。本文提出一种预测感知的自适应回溯框架,用于处理有计划和实时请求的异构多机器人任务分配。问题建模为有限时域随机动态规划,包含机器人-任务兼容性、有序服务要求、路径约束、服务窗口及终时回报要求。所提策略通过采样未来请求场景评估当前分配,同时限制仅对已观测请求做即时承诺。为支持在线应用,框架结合剪枝候选控制、等待动作与交互感知基策略,实现高效未来成本估计。通过根据近期预测误差自适应重加权预测请求,并选择性重新优化已分配但未启动的任务,增强对预测误差的鲁棒性。此外,提出基于历史数据的部署前异构机器人车队配置选择方法。基于医院住院部真实护理任务数据的案例研究显示,该方法实现近完全服务,显著降低服务请求等待时间,优于反应式、令牌传递、预测定位和短视贪心等基线,尾部延迟改善最明显。
原文摘要 · Abstract (English)
Heterogeneous multi-robot service systems must assign requests to compatible robots, construct feasible schedules, and adapt as new tasks arrive online. Historical data can help anticipate future demand, but relying too heavily on inaccurate predictions can degrade performance under distribution shifts. We develop a prediction-aware adaptive rollout framework for heterogeneous multi-robot task assignment with scheduled and real-time requests. The problem is formulated as a finite-horizon stochastic dynamic program incorporating robot-task compatibility, ordered service requirements, routing constraints, service windows, and end-of-horizon return requirements. The proposed policy evaluates current assignments using sampled future request scenarios while restricting immediate commitments to requests already observed. To enable online use, the framework combines pruned candidate controls, wait actions, and an interaction-aware base policy for efficient future-cost estimation. Robustness to forecast error is provided by adaptively reweighting predicted requests based on recent prediction mismatch and selectively re-optimizing assigned but unstarted requests. We also introduce a historical-data-driven procedure for selecting the heterogeneous fleet composition before deployment. In a case study using real nursing-task requests from hospital inpatient floors, the proposed approach achieves near-complete service and reduces serviced-request wait times relative to reactive, token-passing, prediction-positioning, and myopic greedy baselines, with the largest improvements in tail-delay metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。